Wednesday, September 2, 2026

505 papers cs.AI · cs.LG · cs.CL ← 2026-09-012026-09-03 →

Jul Aug Sep

Highlights

UI-Venus-2 Technical Report

Highlight HF pick · 50▲Agents Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao et al. Graphical user interface agents tend to be tuned for benchmarks rather than deployment, held back by narrow environment coverage, brittle task construction, and reward signals that cannot be trusted. UI-Venus-2 scales three axes together for a single closed-loop reasoning-and-action foundation agent spanning mobile, web, and desktop: coverage of more than 170 multilingual mobile apps plus native desktop operating systems, a deep-research pipeline that generates function-grounded instructions, and trace-level plus sample-level verifiers using visual keypoints and multi-model voting to produce reliable reinforcement learning rewards. Safety-aware mechanisms gate consequential actions, and the model is released open source.

Multimodal GUI agents score well on benchmarks but stay brittle in real deployments because environment coverage is narrow, generated tasks are poorly grounded, and reward verifiers are too coarse for reliable reinforcement learning. Ant Group's UI-Venus-2 (9B and 27B, initialized from Qwen3.5-9B and Qwen3.6-27B) attacks all three at once by jointly scaling environments, function-grounded task generation, and multi-model verification, and by unifying mobile, web, and desktop control in a single open-source policy.

  • Training follows a three-stage recipe: large-scale trajectory mid-training, step-level offline RL run separately per domain (Grounding, CAPTCHA, Mobile, Web, Computer), and a multi-teacher on-policy distillation (MOPD) stage that merges the domain experts while conditioning the distillation signal on action correctness, suppressing supervision when the action is right, emphasizing parameter tokens when only they are wrong, and masking parameters when the action type is wrong, with teachers additionally given a hint about the correct action type that the student never sees.
  • Data comes from a closed-loop pipeline that builds per-app capability catalogs via deep research (170+ mobile apps, 4,000+ web domains across 19 categories, desktop tasks serialized as TaskSpec snapshots), and verification uses Semantic Guided Verification, which extracts visual keypoints from the task, judges them window by window with a VLM, and aggregates votes across heterogeneous models into completed/partial/infeasible/failed labels used for data stratification rather than as raw rewards.
  • Headline results for the 27B model: 93.4% on WebVoyager (595-task refresh), 80.2% on REAL, 84.0% on AndroidWorld, 80.5% on OSWorld-Verified, 55.5% on DeskCraft, and 79.9% Pass@1 on the new VenusBench-CAPTCHA (versus 53.0% for Qwen3.6-27B), with the 9B model typically within a few points.
  • Grounding is strong but not uniformly state of the art: 80.1% on VenusBench-GD is a new best, while ScreenSpot-Pro (74.1%) and UI-Vision (66.9%) trail Qwen-UI-Agent-27B, which also leads on MobileWorld at 50 steps (82.1% vs 76.1%).
  • Long-horizon desktop work remains largely unsolved: on OSWorld 2.0 the 27B model reaches only 2.8% binary accuracy and 13.2% partial score against 13.0% for GPT-5.5, and the authors caution that many baseline comparisons use model-specific action scaffolds, different task subsets, and live websites whose state varies by evaluation date.

Safin-1: Safety from Within through Memory-Native State Evolution

Highlight HF pick · 17▲Safety & Alignment Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Zhekai Chen, Cheng Jin et al. Safety in foundation models is usually imposed through external safeguards or post-hoc alignment such as supervised fine-tuning rather than living inside the model's own computation. Safin-1 pursues the alternative through MARCH (Memory-Anchor Routing across Context History), an architecture that maintains structured memory states and retrieves relevant history via content-conditioned routing, supporting test-time adaptation of persistent capability states — including a dedicated Safety State — without repeatedly modifying the backbone. The authors report substantial safety improvements from state-based adaptation alongside evaluations of general capability, long-context understanding, retrieval, and efficiency, and describe the work as an initial architectural exploration rather than a completed programme.

Recurrent and linear-attention models compress the whole prefix into one evolving state, so earlier associations get overwritten and become unrecoverable, and safety behavior is usually bolted on through fine-tuning or external guards. Safin-1 addresses both with MARCH (Memory-Anchor Routing across Context History), which checkpoints the recurrent state into addressable anchors that tokens can route over, and reuses that same routed state bank to host a learned, detachable Safety State so safety is invoked through the model's native computation rather than imposed from outside.

  • Every 512 tokens the recurrent state is snapshotted as a matrix-valued anchor paired with a content-derived 64-dim routing key; each token softmax-routes over causally visible anchors plus a learned zero-payload null option, and the weighted readout is added to the current-state path without altering the base recurrence, with a fused FlashAttention-style reader and Top-4 sparse routing that more than doubles training throughput versus dense routing at 128K tokens.
  • In matched 0.8B pretraining on 50B tokens, MARCH on Gated DeltaNet lifts the 8-task commonsense average from 40.1 to 41.5 (slightly above both full-attention baselines), LongBench from 11.9 to 14.9 (+25%), real-world in-context retrieval from 20.5 to 23.3, and beats the best recurrent baseline in 19 of 24 RULER NIAH settings, with the largest gains on multi-needle tasks and at 32K beyond the 16K training length, and the improvements hold across GDN, KDA, and GDN2 backbones.
  • Scaling via matched continual pretraining and SFT from Qwen3.5-4B and Qwen3.5-35B-A3B raises the ten-benchmark macro average from 66.79 to 69.20 and 76.25 to 78.35, with gains concentrated on hard reasoning (AIME 2025 Avg@64 up 8.0 and 8.2 points, MMLU-Pro +9.3 at 4B, GPQA-Diamond +6.6 at 35B-A3B) while GSM8K, MATH, and MMLU stay within about a point.
  • With the backbone frozen, a persistent per-layer Safety State trained on 1,915 STAR-1/STAR-benign conversations (each replicated with 0, 512, and 2,048 tokens of benign prefix so it works across anchor-bank configurations) cuts average jailbreak attack success rate by 42.3% at 4B and 52.3% at 35B-A3B across WildJailbreak, FORTRESS, StrongREJECT, Jailbreak-R1, and JailbreakBench, with substantially less XSTest over-refusal than a training-matched rank-8 LoRA.
  • The safety-state results table is truncated in the available text, so absolute ASR, over-refusal, and capability-retention numbers cannot be checked, the baseline ASRs are already low (roughly 6-7%) so the relative reductions rest on small denominators, dense routing beats Top-4 on NIAH in the 0.8B ablation even though the scaled models use Top-4, and the authors themselves frame this as only an initial architectural exploration of "Safety from Within."

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

Highlight HF pick · 6▲Multimodal Yijun Chen, Yaqi Zheng, Yanya Li, Boyi Xiao, Buqiang Xu, Shuofei Qiao et al. Memory systems for long-video question answering usually store captions, frames, transcripts, and graph facts as separate fragments, forcing the language model to reassemble cross-modal and temporal alignments at inference time when context is scarce. EM^2Mem instead binds heterogeneous evidence to event anchors while memory is being built, so each event-indexed cell already carries aligned multimodal records, temporal context, graph relations, semantic facts, and provenance. Across three long-video QA benchmarks it gains 2.0, 2.4, and 3.7 accuracy points over the strongest memory baseline and 7.0 points of strict event-level top-5 evidence recall, while cutting per-query latency 4.67 times and total inference tokens by 63.66%.

Long-video question answering needs an external memory, but existing multimodal memories store captions, frames, transcripts, and graph facts as separate fragments that the LLM must re-align across modalities and time at inference, when context is scarce and attribution is hardest. EM^2Mem instead binds all heterogeneous evidence to short event anchors during memory construction, so retrieval returns grounded, generation-ready multimodal events rather than pieces to be stitched together.

  • The video is cut into 30-second segments, each anchored to an event cell holding a keyframe caption, transcript, representative keyframes, and structured fields for actions, objects, topics, scenes, and entities, with multi-scale context views at 3-minute, 10-minute, and 1-hour spans, plus an episodic graph of cross-event entity and temporal links and a semantic graph of recurring habits and relations.
  • At query time a lightweight retriever picks candidate event cells, expands them through graph links, an LLM selector filters them, and the answer model reads a compact evidence view with up to three keyframes for visual verification.
  • Against the reproduced WorldMM baseline under the same setting, average accuracy rises 66.0 vs 64.0 on EgoLifeQA and 76.8 vs 73.1 on Video-MME (L), and 67.7 vs 65.3 on Ego-R1 Bench against published numbers, while per-query latency drops from 459 s to 98 s (4.67x) and total inference tokens fall 63.66%.
  • Strict 30-second event-level Top-5 evidence recall reaches 30.8%, +7.0 points over WorldMM after five retrieval rounds, and ablations show temporal context views matter most (removing them costs 5.6 points), with construction-time unification beating retrieval-time fusion by 3.2 points on structured fields.
  • Margins over the originally published WorldMM numbers are thin (0.4 and 0.2 points on EgoLifeQA and Video-MME (L)), WorldMM stays stronger on several habit and temporal categories, converting visuals to text fields loses fine pixel detail, upstream captioning errors propagate into memory, and the heavy upfront construction cost only breaks even on wall-clock time after roughly 23–24 queries per video.

Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs

Highlight HF pick · 5▲Agents Wentao Zhang, Syed Shariyar Murtaza, Junaid Ahmad Bhatti, Utkarsh Soni, Yifan Nie, Eugene Wen et al. In multi-agent LLM pipelines a single prompt usually carries two entangled jobs — producing task content and specifying execution protocol such as message routing, output format, and termination signals — so an optimizer tuning the content can silently break the protocol and crash the pipeline. The proposed control-data flow separation encodes execution-critical control as typed, validated program objects while leaving only natural-language task content exposed to prompt optimization. Across synthetic reasoning, collaborative review generation, and insurance rating workflows, the framework achieves 100% eventual protocol validity while still improving task performance.

Prompt optimizers such as TextGrad and DSPy treat every line of an agent's prompt as editable, but in multi-agent pipelines those same prompts also encode the routing commands, output formats, and stop signals the surrounding Python controller parses. The paper proposes control-data flow separation: each agent emits a typed, schema-validated control object that only the controller reads, plus a free-form data message that other agents and the optimizer read, so prompt edits can never reach the execution interface.

  • Control schemas are declared as Python dataclasses or Pydantic models with Literal fields for closed sets such as routing targets, the auto-generated JSON scaffolding lives in a frozen prompt slot the optimizer cannot touch, and parse or validation failures trigger bounded retries or a default action rather than ever reaching the router, which the authors formalize as a protocol-stability lemma and ship as the cdsep library.
  • Across four settings (a BBH subset, MARG review generation, and synthetic plus industry-verified insurance underwriting) the method reaches 100% episode stability and the best task score everywhere, including 78.3% BBH accuracy versus 74.3% for DSPy + BootstrapFewShot, 44.4 Jaccard on MARG versus 43.2 for DSPy + MIPROv2, and 36.7% on partner-rated underwriting versus 31.7% for the partner's own hand-written prompt.
  • Naive TextGrad collapses on the routing-heavy tasks, falling to 0% stability on MARG and 56.7% on industry underwriting because the optimizer rewrites inline JSON and chapter-name instructions, and a prompt diff shows it touches control-relevant tokens about four times more often than the separated variant (16.6% versus 4.2% of edited lines on review).
  • Ablations attribute the properties to distinct components: schema scaffolding alone lifts MARG stability from 0% to 100%, parse retry adds only about one point of remaining reliability, and per-example feedback (rather than a scalar batch loss) drives the quality gains, raising review Jaccard from 26.9 to 38.0 and underwriting accuracy from 37.8% to 51.1% at the same budget.
  • The guarantee covers only protocol validity, not semantic correctness, MARG scores rely on a gpt-5.4-mini LLM judge rather than the original GPT-4 setup so are not comparable to the published benchmark, the cross-family runs on Claude and Gemini use a 12-paper subset with a single seed, and the framework assumes fixed agent roles and schemas with no support for runtime agent creation or schema evolution.

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory

Highlight HF pick · 9▲Reasoning Xincheng Wei, Yifan Ding, Yoshua Li, Dongsheng Ma, Rongxiang Weng, Xunliang Cai et al. Self-play lets a model generate its own training questions, but without direction its solver plateaus: existing unguided signals such as difficulty, learnability, or diversity keep questions hard and varied without saying which reasoning weaknesses to attack, while guided methods import direction from human examples or document corpora outside the loop. DiagEvo derives that direction from the solver's own failure history — a diagnostician extracts recurring error causes into a hierarchical memory that groups them under skill nodes and marks each Active or Mastered by self-consistency on targeted questions, and the challenger uses those states and recurrence counts to trade off cause-targeted generation against free exploration, with double-confidence filtering keeping mid-difficulty questions only when the majority answer leads clearly. With a 4B diagnostician it beats every baseline in mean accuracy across all nine benchmarks for Qwen3-4B, Qwen3-8B, and OctoThinker-8B, reaching 72.3% mean accuracy on five mathematical reasoning benchmarks with Qwen3-8B, 4.5 points above R-Zero.

Self-play curricula for reasoning models tend to stall because difficulty and diversity signals say how hard a question is but not why the solver fails. DiagEvo mines the solver's own failed trajectories for recurring error causes, stores them in a hierarchical memory with Active and Mastered states, and lets that memory steer the challenger toward the weaknesses that still need practice, without any external corpus, human examples, or difficulty labels.

  • A frozen Qwen3-4B-Instruct-2507 diagnostician compares a failed GRPO trajectory against a pseudo-label-agreeing one to extract the earliest reasoning difference as a transferable cause, deduplicates it via embedding retrieval, and files it under a skill node; a cause is promoted to Mastered when mean self-consistency on questions targeting it reaches 0.70, and reactivated when new failures match it.
  • The challenger mixes cause-targeted generation (sampling Active causes in proportion to their failure counts, optionally stitched with a Mastered sibling from the same skill node) with free exploration, with the mixing odds scaling linearly with the normalized failure count so that mastered causes push probability back toward exploration.
  • Double-confidence filtering keeps only questions whose majority vote share lies in 0.25 to 0.75 and whose top answer leads the runner-up by a ratio of at least 1.6, which raises oracle agreement of retained pseudo-labels from 65% to 75% at round 5 versus unguided self-play.
  • On Qwen3-8B-Base it reaches 72.3% mean over five math benchmarks and 57.4% over all nine, 4.5 points above R-Zero and 1.1 above DARC, while also topping every baseline on Qwen3-4B and OctoThinker-8B; scaling the diagnostician to 235B adds only about 1 point, and the added diagnosis machinery costs 5.7% of wall-clock time.
  • Gains still plateau and reverse after round 5 (72.3% falls to 71.4% by round 7) with the round count fixed in advance, the filter measures agreement rather than correctness so shared solver errors can pass, and the curriculum is math-only with general-reasoning gains reported as transfer rather than direct construction.

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Highlight HF pick · 55▲Large Language Models Shaowen Wang, Ge Zhang, Kairong Luo, Yuhao Wu, Shaofan Liu, Jiaheng Liu et al. Looped transformers add effective depth by iterating a shared block, but comparing at fixed model size hands the looped variant extra floating-point operations, conflating architecture with compute. SMELT loops the middle half of layers twice in a Mixture-of-Experts transformer while matching per-token FLOPs, total non-embedding parameters, and key-value cache against an unlooped baseline, scaled across four sizes up to 54B non-embedding parameters with a separate Chinchilla-style scaling law fit per architecture. Loss falls faster with compute, saving 6.8 to 18.0 percent of training FLOPs on the compute-optimal frontier, with downstream gains exceeding what validation loss predicts, largest on code, and growing with sequence length and in-context example count. Mechanistic analysis credits the second visit with shrinking the attention sink and redirecting attention toward content-relevant tokens.

Looped Transformers add effective depth by re-running a shared block of layers, but prior comparisons held parameter count fixed and let the looped model spend extra FLOPs and KV cache, so the architectural benefit was never isolated. The authors use Mixture-of-Experts to match per-token FLOPs, total non-embedding parameters, and KV cache simultaneously, then ask whether looping still wins across a full scaling ladder.

  • The SMELT recipe loops the middle 50% of layers twice, pays for the second pass by narrowing the hidden dimension, restores capacity by raising expert count (192 to 288 at the 200M scale), scales looped residual updates by 1/2, and shrinks head size at a higher GQA ratio so all three budgets land within about 4%; ablations at 200M showed a 50% span beats full-stack looping, two visits beat three or four (which force a thinner model), and the looped model tolerates a larger effective depth-to-width ratio than the unlooped Baseline.
  • Across a 4×4 grid of scales (100M to 1.6B active, up to 54B non-embedding parameters) and compute-equivalent sparsities of 85/95/97%, separately fitted Chinchilla-style surfaces give SMELT a steeper frontier exponent (γ = 0.250 vs 0.237), translating to 6.8–10.0% training-FLOP savings at 10²⁰ FLOPs and 14.7–18.0% at 10²¹, while compute-optimal tokens-per-parameter stay within 6% of the Baseline, so the saving comes from lower loss at the same allocation rather than reallocating budget.
  • Downstream, SMELT wins 96/96 matched pairs on DCLM Completion, 83/96 on DCLM Core, and 29/30 above-chance pairs on MMLU, and its residual above a sigmoid calibration from validation loss to benchmark score is positive at every scale and grows with model size, with Code the most-improved training domain and gains that widen with sample length and number of in-context examples.
  • Mechanistic probes show the second visit selects largely the same experts and attends largely the same tokens as the first but produces larger, aligned residual updates that amplify rather than overwrite, changes values more than queries and keys, and reduces attention-sink mass in favor of content-relevant tokens.
  • Caveats: every configuration was trained once so reported error bars exclude training-run variance, the cell-bootstrap intervals are wide (10²⁰ FLOPs at 85% sparsity spans [1, 22]%), the 10²² FLOP gains of 19.6–23.5% are extrapolated beyond the fitted window with intervals reaching zero, the dense S=0 control was excluded from the fit, and the data and Baseline family are internal so the loop-span choice rested on validation loss because DCLM metrics did not track it in the ablation.

Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching

Highlight HF pick · 17▲Robotics Jaewoo Park, Minyoung Lee, Sukmin Seo, Moonbin Yim, Hyunwook Yoon, Dohoon Ryu et al. How far multimodal large language models (MLLMs) extend from perceiving to acting is tested by dropping one directly into a drone's control loop with its entire action space declared solely in the prompt. DroneCATS-Agent makes the model a swappable component and the DroneCATS benchmark treats it as the independent variable across approaching a visible target, tracking a moving one, searching outside the initial view, and commanding a multi-drone fleet, with no fine-tuning or function-calling schemas and a roster scaling down to 2B parameters. Small open models often navigate into the success radius more reliably than frontier models yet lose the episode by declaring arrival prematurely or never at all, and multi-drone commanding widens the gap as small models blindly copy one coordinate across distinct views, making protocol discipline and correct termination, not navigation, the bottleneck.

Prior MLLM drone-control systems (TypeFly, SPF, Fly0, OnFly) progressively strip decisions away from the model, which is why backbone choice appears not to matter; this work holds a minimal agent fixed and makes the model the independent variable. DroneCATS-Agent declares four actions (go, rotate, think, finished) purely in the prompt with no fine-tuning or function-calling schema, and the DroneCATS benchmark scores approaching, searching, tracking, search-and-track, and four-drone commanding under one criterion anchored on the model's own arrival declaration.

  • Each step the model sees only an egocentric monocular RGB frame, the instruction, and its last five actions, and returns one JSON action: go gives a pixel plus a self-estimated depth that a rule-based controller lifts to a velocity setpoint, rotate yaws in place, think hovers and unlocks reasoning for the next call, and finished is a claim the scaffold logs while the episode keeps flying, judged post hoc as success only if some declaration occurred within 5 m of the target while it was visible (100 AirSim episodes, 300 s cap).
  • Even the simplest cell is unsolved: the best model, Gemini 3.7 Flash, hits 65% on approaching, 40% on searching, 80% on tracking, and 45% on search-and-track; no model exceeds 40% when the target is withheld from the first frame; GPT-5 approaches at 60% but tracks at only 15%; the embodiment-specialised Gemini Robotics-ER 2 averages 47.5% versus its generalist sibling's 57.5%; and the Qwen3.5 ladder falls monotonically 33.8/21.3/12.5/0% from 27B to 2B.
  • The headline finding is that small models fail on protocol, not navigation: Qwen3.5-9B enters the success radius in 90% of approaching episodes, more than any frontier model, yet converts only 35% by declaring at 0.63 of its start distance, while Qwen3.5-2B declares at 1.28× its start distance and never succeeds, and Cosmos3-Edge-2B reaches the radius in 25% of episodes but never declares at all.
  • Commanding four drones from one context toward look-alike targets distinguishable only by a licence plate or sign amplifies this split: Gemini 3.7 Flash wins 80% of episodes and Gemini Robotics-ER 2 65%, but GPT-5 drops to 20%, Claude Opus 5 to 0%, and Qwen3.5-9B and Qwen3.5-27B paste an identical point across all four distinct views on 70% and 58% of go-steps, so at most one command can be grounded.
  • Caveats are stated plainly: 20 episodes per cell yields a per-cell standard deviation of 9–13 points (a three-flight audit gave 54.6 ± 9.7% overall), so orderings within a tier are within noise; results are simulation-only; success degrades when the control loop slows below ~2 Hz; the think action was invoked 1,797 times but measurably changed nothing (+0.03 m/step, reacquisition 13.5% vs 13.4%); and hierarchical commander-plus-subagent control is left unevaluated.

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

Highlight HF pick · 3▲Agents Haoyang Yan, Min-le Su, Hangfan Zhang, Zhanhao Li, Chen Zhang, Shao Zhang et al. Autonomous software development asks LLM coding agents to turn high-level requirements into complete working systems without human intervention. Harness-of-Harness (HoH) wraps existing coding-agent harnesses in iterative planning-coding-testing loops, balancing repair against capability growth, scoping work into small verifiable increments, separating implementation-time testing from independent evaluation, progressively exposing deliverables, role-specific tools, and skills, and maintaining versioned project history. Across GameCraft-Bench, FrontierSWE, and ProgramBench with three harness-model pairs (Codex with GPT-5.5, OpenCode with DeepSeek-V4-Pro, and Pi with MiniMax-M3), it beats the standalone harnesses by an average relative gain of 52.25 percent after three iterations, peaking at 82.86 percent. A multi-day deployment of more than 70 iterations produced a playable first-person-shooter game with a coherent storyline, implemented core mechanics, visuals, and audio.

Coding agents work well within a bounded episode but struggle to carry a from-scratch software project forward over many sessions, losing track of requirements, unresolved failures, and already-verified behavior. Harness-of-Harness (HoH) wraps an unmodified coding harness in a repeated planner–developer–tester loop that hands both the evolving artifact and structured test evidence from one iteration to the next.

  • Each loop invokes the same harness–model pair three times in fixed roles: a Project Planner reads the spec plus prior test evidence and writes a bounded development document, a Developer (the sole writer) implements it with shift-left self-tests, and a read-only QA Tester evaluates a frozen candidate with black-box and white-box checks and returns a structured report; outputs that violate the required schema trigger a retry, while reasoning and tool use are left unconstrained.
  • After three iterations, HoH beats the standalone harness on all three benchmarks for all three configurations (Codex+GPT-5.5, OpenCode+DeepSeek-V4-Pro, Pi+MiniMax-M3), with an average relative gain of 52.25% and absolute gains of 16.62–22.08 points on GameCraft-Bench, 19–29 points of dominance on FrontierSWE, and 6.09–16.85 points of test pass rate on ProgramBench; extending Codex to ten loops on FrontierSWE lifts dominance from 39.33% at HoH@3 to 72.67%.
  • The gain is not just extra compute: at matched pass counts on GameCraft-Bench, HoH outscores repeated-session Vanilla Continuation by 10–13 points, and HoH@2 at 5.67M tokens (64.84) beats three-pass continuation at 6.33M tokens (58.24); ablations show freezing the plan, dropping test evidence, or rebuilding from scratch each loop costs 6.28–8.13 points.
  • In a multi-day run with Codex and GPT-5.6-Sol, plus Godot MCP tooling and asset, UI, and testing skills, HoH built a playable narrative first-person shooter over 70 loops, closing 65 of 81 recorded issues, though 16 remained open and 17 issues were reopened after later changes regressed verified behavior.
  • Benchmarks were sampled (45 of 140 GameCraft-Bench tasks, 15 of 17 FrontierSWE tasks) and the multi-day case is a single project with one configuration, so generalization beyond game development and the tested harnesses rests on limited evidence, and Pi peaked at HoH@2 rather than HoH@3 on ProgramBench.

H3-World: Turning Language Understanding into World Control

Highlight HF pick · 30▲Multimodal Danze Chen, Zeqing Wang, Ziyue Lin, Xingyi Yang, Yeying Jin H3-World turns the 33B MiniMax-H3 video generator into an interactive world model by exploiting the observation that large video generators already accept coarse natural-language control of character behavior and camera motion. Each action is represented as a structured combination of character and camera instructions aligned to the corresponding temporal video latents, and temporal attention routing confines each instruction to its intended time interval to reduce control leakage across actions, with no dedicated action modules added. Using only 8,000 gameplay samples, 10,000 LoRA optimization steps, and 0.199% trainable parameters, the system achieves effective character and camera control, preserves generation quality, and generalizes to unseen scenarios.

Interactive world models normally need action-conditioned video generators built with dedicated control modules trained on aligned action-video trajectories. H3-World instead observes that the 33B MiniMax-H3 video generator already follows coarse language instructions for character and camera motion, and turns that into precise, temporally grounded control by expressing each action as text bound to a specific video latent interval and adapting only 0.199% of the parameters with LoRA.

  • Each discrete control state (eight character and camera keys plus a camera-speed flag) is mapped to a short compositional instruction such as "the man walks backward and strafes left, camera pans right slowly," drawn from 9 character clauses and 16 camera clauses, and injected through the model's native text pathway rather than a learned action embedding.
  • Every video latent interval gets its own independently encoded action prompt with a mirrored temporal position, and a deterministic single-egress attention mask lets each action span be read only by its matched latent, so actions enter the video stream at one point and then propagate through the unchanged bidirectional video-to-video attention.
  • Training uses only 7,872 gameplay clips from ABot-World-Explorer-500h (124 frames, 832×480, 37 action prompts per clip) with rank-32 LoRA for 10,000 steps on the attention projections and a two-layer token refiner, keeping the backbone, encoder, and VAE frozen.
  • On a controlled left-then-right pan schedule, H3-World produces cumulative horizontal flow of +52.7 before and -106.0 after the switch, while frozen H3 with a global prompt gives 0.0 and -17.3 and a zero-LoRA per-latent interface stays nearly static, showing that both the temporal binding and the adaptation are needed; text-based control also beats additive-bias and FiLM action conditioning in qualitative comparisons.
  • The model composes unseen character-camera pairs (52 of 135 valid combinations never appear in training) and transfers to out-of-distribution scenes, but evaluation is mostly qualitative on representative examples, generation is fixed-length and short-horizon, and there is no persistent world state, real-time interaction, or planning.

StudentSim: Training LLM-based Student Simulators

Highlight HF pick · 248▲Applications Ke Yang, Chenglong Wang, Michel Galley, Chandan Singh, Jeevana Priya Inala, ChengXiang Zhai et al. AI tutors work best when they adapt to individual students, but evidence about which guidance suits which learner is slow and costly to gather, and existing student simulators either track state without processing explanations or role-play fluently without matching the target student's competence. StudentSim turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization, so each simulator both mirrors a student's own responses and updates them under tutor guidance. The accompanying StudentSimEval protocol covers 60 students across chess, second-language English writing, and mathematics and measures behavioral fidelity and guidance responsiveness; StudentSim outperforms GPT-5.4 on both metrics in all three domains, reaching fidelity 0.51 and responsiveness 0.91 in chess versus 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. Used as a reward model for tutor reinforcement learning, it produced a chess tutor that expert humans rated more accurate, better-guided, and more personalized than both a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward.

Real-student feedback on which tutoring moves work for which learner is sparse and slow to collect, and existing simulators cover only half the job: behavior-tracking models mimic a student but cannot read tutor guidance, while prompted LLMs follow guidance but fail to reproduce a specific student's competence. StudentSim trains one small simulator per student by first pooling records across all students in a domain and then specializing that base on each individual's sparse data, and StudentSimEval scores any simulator on both behavioral fidelity and guidance responsiveness.

  • Both stages fine-tune LoRA adapters on Qwen3-4B-Instruct, mixing single-turn records that teach a student's own responses with multi-turn records that pair a wrong answer, tutor guidance, and the canonical correction, so the model learns to update toward the fix rather than just imitate.
  • The benchmark covers 60 students across chess (Lichess games), second-language English writing (EFCAMDAT), and multiple-choice math, and evaluates every method on the same frozen per-student held-out records.
  • In chess, StudentSim reaches fidelity 0.51 and responsiveness 0.91, versus 0.23 and 0.72 for prompted GPT-5.4 and 0.45 and 0.27 for Maia2, with the same ranking holding in L2 writing (0.56 / 0.64) and math (0.64 / 0.92); the gap is largest on Socratic and conceptual guidance, where the tutor never states the answer.
  • As a proof of concept, using the pooled simulator as a GRPO reward for a Qwen3-VL-8B chess tutor yields guidance that expert players rate at 90.5% factual accuracy versus 75.7% for the no-RL baseline and 71.6% for a tutor trained against a GPT-5.4 simulator reward, with higher guidance and personalization scores as well.
  • Limitations include very small per-student evaluation sets outside chess (tens of records per learner), LLM-generated rather than real tutor guidance for chess and math, a tutor-RL demonstration confined to chess because it needs an engine-based reward, and the metrics capturing only a single-step update rather than learning, retention, or forgetting over time.

Large Language Models 105

InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information

Jiaze Li, Aocheng Shen, Bing Liu, Boyu Zhang, Xiaoxuan Fan, Qiankun Zhang et al. cross-listed Competitive-programming benchmarks for large language models (LLMs) hand over the full problem input up front, which skips the ability to reason when facts are revealed only in response to queries. InteractBench gathers 322 interactive problems from Codeforces, AtCoder, IOI, and ICPC, each packaged with an executable local interactor so generated programs can be judged fully offline across multi-round exchanges bound by protocol rules and query budgets. Even the strongest reasoning models achieve only limited success on these problems, and a fine-grained failure taxonomy traces the gap mainly to algorithmic logic errors but also to frequent protocol violations and exhausted query budgets.

Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning

Yuxuan Li, Victor Zhong, Ehsan Kamalloo Persona prompting usually relies on synthetic personas that flatten real variation and lean on stereotypes rather than the signals that actually drive preferences. Profile behavioral grounding instead derives open-ended, high-fidelity user profiles from authentic anonymized social media posts, then tests them both as training data for supervised fine-tuning and as non-parametric context for test-time multi-perspective reasoning. Behavior-derived profiles beat synthetic-persona baselines and improve base models on recommendation and open-ended query benchmarks under both paradigms, with code released publicly.

ES-AHD: An Evolution Strategy Framework for Automatic Heuristic Design

Yutao Lai, Kezhao Lai, Hai-Lin Liu, Yuping Wang, Ping Guo cross-listed LLM-driven automatic heuristic design typically mutates individual programs at random, which searches blindly and balances exploration against exploitation poorly. ES-AHD borrows two ideas from evolution strategies: semantic recombination replaces point-to-point reproduction by having the LLM extract shared insights from top-performing individuals to define a search direction, and the covariance matrix is mapped onto the model's sampling temperature so the search radius shrinks for local code refinement while occasionally spiking to escape semantic local optima. The result is a directional, center-guided sampling scheme that the authors report accelerates the discovery of high-quality heuristics, with source code released.

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

Qian Zhang, Yaoming Li, Zhewen Tan, Yanshu Wang, Heng Lu, Kun Su et al. Post-training quantization methods solve each layer with one closed-form second-order step, which forces heavy approximation of the global loss — dropping cross-channel coupling, pooling output rows — and then freezes the resulting Hessian for the whole layer even as the loss landscape shifts column by column. REAL-Q keeps an end-to-end-aligned surrogate of the global loss and refines it with block-wise gradient descent after every 128-column block, adding a sliding window across layer boundaries to limit error propagation. On LLaMA-3.1 at 8B and 70B and Qwen3 from 0.6B to 32B at 4-bit weights with 16-bit activations, it cuts end-to-end KL divergence by up to roughly 49% against state-of-the-art globally guided methods.

Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning

Jinyuan Zhang, Peng He, He Hu, Yin Yuan, ShengShuo Jiao Fine-tuning can erode in-context learning (ICL), and a common diagnostic assumes a model is still context-sensitive if its attention shifts when the demonstrations change. Formalizing that proxy as In-Context Sensitivity (ICS) — the average row distance between last-token attention on matched versus mismatched demonstration prefixes — and pairing it with the behavioural accuracy gap, a four-arm ablation on Llama-2-7B shows an ICS-maximising regulariser pushing ICS to 1.413, within 0.5% of its geometric ceiling, while the behavioural gap stays near zero and MMLU accuracy falls from 0.371 to 0.279. Endpoint analysis finds attention becoming sharp and near-disjoint across prefixes but routing to formatting and demonstration-body tokens rather than labels, a textbook Goodhart failure; objectives anchored to the pretrained computation instead hold a high-MMLU, moderate-ICS region.

OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization

Yishan Yao, Binjun Li, Hanling Yi, Pengyu Li, Xiaoqing Liu, Zihan Yang et al. NVFP4 is a microscaling format for low-bit inference, but a single large activation in a quantization block dominates the shared block scale and inflates the error of every other value in that block — an effect the authors name Collateral Quantization Error. OCGQuant attacks it purely through channel grouping, adaptively pairing outlier channels with low-magnitude companion channels so blocks are composed more favourably, without the mixed precision, rotations, or residual compensation that other post-training quantization methods add. On Llama3 and Qwen3 it achieves the lowest WikiText-2 perplexity and the highest average downstream accuracy among the evaluated methods while keeping prefill speedup close to round-to-nearest and matching its peak decoding memory.

KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training

Meghanadh Pulivarthi, Kushagra Bhushan, Vineet Kumar, Gaurav Pandey, Jaydeep Sen, Dinesh Raghu et al. Continued pre-training struggles to install knowledge from niche documents such as manuals or technical specifications, because those sources rarely repeat a fact, and the usual remedy of generating many paraphrases requires an expensive teacher model. KItCAT (Knowledge Injection via Corrupted Auto-regressive Training) augments ordinary next-token prediction by randomly replacing a subset of input tokens with other vocabulary tokens while leaving the next-token labels unchanged, producing many varied training inputs per document at negligible cost. The corrupted-input training consistently outperforms standard continued pre-training across multiple datasets and model families, reducing the need for paraphrase generation in decoder-only models.

Commit-first LLM judging inherits the judge's own errors

Idil Gozel cross-listed One defence against systems gaming an LLM judge is commit-first judging, where the judge solves the task itself, commits to an answer, and accepts a candidate only if the two match. An audit of default judge configurations in eight widely used evaluation frameworks finds none of the 24 configurations in scope implement it, while nine use a variant the literature measures as ineffective, all traceable to a single ancestor prompt through a copied typographical error. In a controlled experiment, plain best-of-N search with no access to correct answers got 90 of 96 candidates accepted on an interval-merging task under one documented configuration, with every accepted program failing a held-out suite; commit-first judging dropped acceptance to zero there, but on a second task the judge's own committed answer was wrong and the population converged on it, showing the defence relocates the gameable anchor to the judge rather than eliminating it.

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

Zhigeng Liu, Zhiyuan Ning, Ruixiao Li, Xiaoran Liu, Yuerong Song, Min Zhang et al. Long-context decoding in large language models is bounded by memory bandwidth and the quadratic cost of attention, and existing sparse-attention schemes trade the memory overhead of metadata indices against the compute cost of adaptive block selection. Faster Flash Decoding (FFD) fuses the selector and the computer into a single kernel, replacing external metadata with content-aware scanning under low-bit quantization, and adds a top-delta rule that filters blocks to a distribution-adaptive sparsity level without global synchronization. The training-free, drop-in design reuses scanning results for the subsequent computation and reaches up to 11.6x kernel-level speedup and 2.37x end-to-end throughput at context lengths up to 256K, with accuracy preserved on RULER and LongBench.

Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs

Deniz Bayazit, Badr AlKhamissi, Antoine Bosselut Claims that multilingual language models route computation through language-specific states such as an English pivot rest on probes that infer a latent language from different signals, either hidden state geometry or what can be decoded from intermediate representations. Comparing these probes across model families, training regimes, domains, tasks, checkpoints, and up to 27 languages shows they systematically disagree: the Gaussian mixture model representation probe finds cross-lingual mixing earlier in the network, while decoding-based probes retain sharper, more English-biased language-specific signals. The disagreement tracks how multilingual a model is and how far training has progressed while staying comparatively stable across domains, which the authors read as evidence that current probes expose different aspects of multilingual processing rather than a single internal lingua franca.

Do General NLP Embeddings Capture Ontological Reasoning?

Hamed Babaei Giglou, Jennifer D'Souza, S\"oren Auer Whether general-purpose text embedding models encode logic-sensitive ontological structure is tested with AVA, a benchmark of 171,007 contrastive triplets built from 163 heterogeneous ontologies via hierarchy inversion, relation substitution, and disjointness injection, each pairing a statement with an equivalent paraphrase and a hard negative whose relational meaning contradicts it. Across more than 25 current embedding models, the best reaches 0.739 triplet accuracy but only 0.135 on the hard negatives. Fine-tuning improves discrimination substantially yet transfers poorly to downstream Semantic Web tasks such as taxonomy discovery and ontology alignment, and further analysis suggests much of the gain comes from recognizing the specific perturbation patterns rather than from robust ontological understanding.

Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs

Jonathan Zheng, Zirui Shao, Alan Ritter, Wei Xu Because language models are trained on static corpora, their knowledge goes stale, and existing tests of knowledge editing either get contaminated quickly or rely on counterfactual edits that clash with facts the model already holds firmly. The authors build ParallelEvents, a benchmark of fictional but plausible future worlds that generates coherent event trajectories, sidestepping contamination while staying internally consistent, and pair it with Synapse, a training framework that updates parameters using model-generated data through mid-training and instruction tuning. Synapse beats existing knowledge-insertion methods by 14.23%, suggesting simulated synthetic data can integrate new knowledge at scale without hand-curated corpora.

LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts

Daniela Occhipinti, Andrea Piergentili, Marco Guerini Sociodemographic prompting conditions a language model judge on an annotator's demographic profile in hopes of reproducing that group's judgments on subjective tasks. Comparing the predicted label distributions of 23 open-weight models against real annotator groups on three tasks, under no demographic information, single-attribute profiles, and intersectional profiles over gender, age, race, and education, the authors find that an unconditioned judge is not neutral — it best matches White, college-educated annotators — and that demographic conditioning is asymmetric, moving predictions toward majority groups and away from minority groups, most sharply on offensiveness where intersectional profiles amplify the effect. Comparing base against instruct models points to instruction tuning as a likely source of the asymmetry, suggesting the technique often fails the minority groups it is invoked to represent.

QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization

Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi Weight-only post-training quantization cuts the cost of serving large language models but degrades badly below 2 bits, and the usual fix — unstructured sparsity — sacrifices the regularity that makes GPU kernels fast. QTEA quantizes weights to ternary values and compensates with salient weights as residual error correctors, confining those residuals to selected columns under semi-structured 1:4 sparsity, and adds column-wise rescale refinement plus an error-decay term that counters the order-dependent error accumulation in GPTQ-style column-by-column quantization. On Qwen3-14B it compresses all weights to an effective 1.7 bits per weight while improving average accuracy 16.7% over the strongest ternary baseline, with similar gains on Llama3-8B, and a lookup-table kernel delivers 7.2× faster per-token generation than FP16.

Authority Bias in Conversational Search Engines for Academic Paper Recommendation

Uthman Jinadu, Parsa Ghazvinian, Anjila Budathoki, Benjamin M. Ampel, Rajshekhar Sunderraman, Yi Ding Whether large language models recommending academic papers judge on content or on prestige signals had not been tested causally. Holding title and abstract fixed, the authors varied author prestige, venue, and citation metadata across original, flipped, and boosted conditions for eight models (five open-weight, three closed) in single-turn top-1 recommendation. Authority bias proved substantial and directional, varying widely by model and only partly reducible through prompt-level debiasing. They also document a say-do gap: debiasing instructions suppress explicit mentions of authority much faster than they suppress authority-driven choice flips, so audits based on what a model says systematically understate its behavior.

Hypotheses-Guided Self Distillation for Continual Personalization

EunJeong Hwang, Kushan Mitra, Dan Zhang, Hannah Kim, Estevam Hruschka Personalizing an assistant over long-term use is hard because users rarely state preferences outright; they leak through heterogeneous, latent, noisy signals, and current approaches either stuff raw histories into context or run expensive reward-based optimization. HypReflect instead infers explicit, uncertainty-aware preference hypotheses from those signals, reflectively revises them as evidence accumulates, and folds the resulting user model back into the assistant through hypotheses-guided self-distillation. It beats raw-history and incremental-update baselines across online personalization, multi-session interaction, and implicit behavioral signals, and generalizes to unseen users and new domains while staying stable across context budgets.

Latent Mechanisms of Language Control in Multilingual Language Models

Ryo Mitsuhashi, Sabri Boughorbel, Majd Hawasly Multilingual models sometimes switch languages mid-generation without reason, which suggests identifiable internal features that control output language. Three ways of locating such features in cross-layer transcoders are compared: selection by activation value (ValSel), by activation frequency (FreqSel), and by LLM-generated annotations of what each latent does (AnnSel), evaluated on two new code-switching benchmarks spanning seven languages with intervention experiments on Gemma-2-2B and Qwen3-4B. All three steer generation language effectively, with FreqSel strongest overall and AnnSel offering human-readable selection criteria. A knock-out analysis shows the three methods pick largely non-overlapping yet individually sufficient latent subsets, pointing to redundancy rather than one canonical language direction.

The Curse of Multilinguality in Lexical Normalization

Saman Rahbar Rewriting non-standard text like tmrw or gr8 into standard forms suffers from scarce labelled data, so practitioners often train one model across many languages at once. Holding a character-level model at fixed capacity and sweeping the number of jointly trained languages from one to twelve on a standard benchmark, per-language accuracy peaks when a language is paired with only one to four others and then declines by roughly forty percent as the remaining languages are added. A control that keeps total training data constant makes the drop arrive earlier and go deeper, implicating capacity competition rather than data volume, and typological distance from the other languages gives no dependable rule for how many co-training languages are best.

Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance

Teng-Ruei Chen Conformance suites for quantized matrix-multiply kernels check whether two implementations agree within a numerical tolerance, and this work measures what such a check can actually catch. Injecting nine faults into a reference INT8 pipeline across 8,232 layer-fault-regime cells of Qwen3-1.7B, all five epilogue faults, including scale precision, double rounding, multiplication order, output truncation, and fused ordering, move the output by at most one bfloat16 spacing, so a one-spacing tolerance is blind to the entire class by construction and misses four of the five outright. Faults that break the accumulator's exactness preconditions or operand sharing are always caught, so the suite establishes those properties rather than interchangeability. Requantizing every weight scale to the nearest power of two makes CUTLASS and Triton agree bitwise at every linear layer (196/196 and 252/252, versus 8/196 and 10/252 with the original scales) and produces byte-identical token sequences at 1.7B, 8B, and 14B, with perplexity shifts under 1% and a previously reported +157% regression traced almost entirely to a probe that rewrote scales without requantizing weights.

Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning

Vishvesh Bhat Data engineering uses of language models, such as translating natural language into SQL and automating spreadsheet operations, need both better accuracy without finetuning and relief from the quadratic cost of attention over long contexts. The authors propose a drop-in neurosymbolic layer that slots into existing model backbones to strengthen logical reasoning and compress context, reporting an average accuracy increase of 85% across benchmarks including BIRD-CRITIC and LiveSQLBench with no task-specific finetuning or RLHF. Applied to long-context inference, the same symbolic prioritization and compression cuts effective token usage by more than half and brings effective time complexity down from quadratic to roughly linear on certain long-context tasks.

Removable and Irreducible: A Token-Cost Ledger for the Multilingual Tokenization Tax

Madhulatha Mandarapu, Sandeep Kunkunuru Non-English text costs several times more tokens for the same content, and because attention is quadratic in sequence length that inflates compute as well — but how much of the penalty is actually removable is unclear. Framing tokenization as source coding with a Shannon-rate floor, the authors assemble a token-cost ledger that splits each language's cost at fixed parallel content into removable coding redundancy, residual coding slack, an intrinsic-content term, and an irreducible grapheme-to-phoneme term. On FLORES-200 across eight languages, a production tokenizer charges up to 8.9x more tokens for Indic scripts than for English, yet a script-matched code trained on only 1,012 sentences removes a median 64% of that excess, and a script-fair information floor shows intrinsic content differs by under 6% — the tax is representational rather than informational, and implies up to 79x attention cost. The authors scope this explicitly as compute-and-memory accounting rather than a model-quality claim, and release a one-command harness.

DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference

Xiaoyang Lu, Belthangady Akash Vi Narayana Pai, Xian-He Sun cross-listed Mixture-of-Experts (MoE) models scale large language model inference cheaply in arithmetic but move a lot of weight data, which becomes the bottleneck on neural processing unit (NPU) systems; near-data processing (NDP) can absorb some of it, yet existing NPU-NDP MoE systems ignore hardware heterogeneity, dynamic expert-level concurrency, and temporal expert reuse across a batch. DynaNDE adds an analytical performance model covering heterogeneous hardware, data-movement cost, and communication-computation overlap, uses it to decide per layer which experts run on the NPU versus near memory while accounting for expert concurrency, and pairs that with a reuse-aware runtime that skips reloading experts already resident in NPU memory. Against the state-of-the-art NPU-NDP MoE serving framework it reports average speedups of 2.6x for prefill and 2.2x for decoding.

Late Transformer Layers Recode Syntax Canonically: Evidence from Greek Scrambling and Cross-Layer Generalisation

Christos Nikolaos Zacharopoulos, Revekka Kyriakoglou, Chara Tsoukala, Th\'eo Desbordes Probing work has established that syntactic information is decodable from early and middle transformer layers, but what happens to it in later layers is poorly understood. Cross-layer generalisation analysis on three Greek-tuned large language models, using minimal pairs of Modern Greek object-relative constructions that differ only in canonical subject-verb-object versus non-canonical verb-subject-object order while preserving meaning, finds that a probe trained on layers 20-31 transfers below chance to each early layer individually, classifying 99.3% of non-canonical sentences as canonical. Probe coefficients reverse sign around layer 22, which the authors interpret as a directional recoding toward the canonical form rather than simple information loss. The result characterises a representational format change beyond the known decline in syntactic decodability, and yields a directly testable prediction for human EEG and MEG decoding on the same stimuli.

HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference

Chun-Ting Chen, Dongmin Han, Hangyeol Mun, Jake Hyun, Arnab Raha, Amit Agarwal et al. Block quantization quantizes both weights and activations of large language models so inference runs on one low-precision datapath, but its design space of bit-width, block size, scaling scheme, and numeric format is largely unmapped. A design-space exploration shows larger blocks amortize dequantization and accumulation cost yet hurt accuracy, motivating HBQ (Hierarchical Block Quantization), which keeps large blocks for efficiency and adds a cheap significand-based second-level scaling that compensates block error better than power-of-two or integer scaling. The accurate variant reaches W4A16-level accuracy at W4A5 in less silicon area than NVFP4, and a 28nm ASIC applying HBQ to weights, activations, and the KV cache delivers 2.3x area and 4.6x energy efficiency over state-of-the-art weight-only quantization at matched accuracy, plus 1.5-3.0x speedup over prior block quantization.

Can LLMs Use Relational Transformer Embeddings?

Francisco Galuppo Azevedo, Clarissa Lima Loures Feeding frozen relational-encoder embeddings into a language model as soft tokens looks like an elegant division of labor, letting the encoder handle multi-table structure while the LLM handles reasoning without lossy text serialization of the database. The test injects embeddings from a frozen Relational Transformer into Qwen3.5-4B through a learned MLP projection plus LoRA, trained with supervised fine-tuning on chain-of-thought traces and then group-based reinforcement learning (GSPO), evaluated on 10 binary classification tasks over 6 RelBench databases under single-task, within-dataset, cross-dataset, and all-task supervision. The hybrid fails to consistently beat the standalone relational encoder, is frequently below random, and proves highly sensitive to serialization format and relational-token budget as well as unstable under RL; the authors publish the negative result and argue soft-token fusion needs stronger alignment objectives and schema-aware design.

Toppling the Hierarchy in Byte-level Language Modeling

Lukas Edman, Alexander Fraser Byte-level language models still fail at seemingly trivial character manipulation, and the leading designs are hierarchical, starting from bytes, downsampling to word-like units, then upsampling back to bytes for efficiency. Comparing hierarchical variants against pure byte-level models shows the hierarchy itself is what limits character-level understanding, with flat byte models consistently winning on character manipulation tasks. Ablating transformer layers into attention and feed-forward parts localizes the effect to byte-level attention as the primary mechanism, establishing an explicit trade-off between computational efficiency and fine-grained character understanding.

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

Qiaoyuan Zheng, Yiqu Yang A gap of under a percentage point on a leaderboard is routinely read as one model being better than another, but that ordering may be an artifact of which benchmark items happen to be included. Using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item response theory, the authors identify items with low residual differential item functioning across model families in one owner-disjoint fold, then use those frozen, difficulty-balanced weights to rescore models in the other fold, with length-matched random subtests as a control. Overall rankings barely move (Kendall's tau-b of .900 to .948), yet in four of five benchmarks 30.9-47.1% of cross-family pairs initially within one percentage point flip order, 16.9-28.6 points above the matched-random baseline, so sub-one-point leaderboard gaps need explicit evidence that the ordering survives benchmark recomposition.

Human-Anchored Factuality Evaluation with Strategic Annotation

Yu Wang, Craig Erickson, Kevin Small Automated factuality judges scale evaluation but drift systematically from human ratings, so the setting studied is a hybrid one: run the judge over the whole dataset, buy human labels for a small subset, and combine them into a statistically valid estimate under a fixed annotation budget. The core observation is that judge-human disagreement in factuality is not just low confidence but has structure — incomplete evidence, temporal mismatch, unverifiable claims, rubric misalignment — so the authors build an annotation policy from failure-space analysis (FSA) signals that predict where the judge and humans will diverge. Against uniform and uncertainty-driven sampling, the FSA-guided policy raises effective sample size by 40.3% on an internal reference-based system (AutoFA) and 27.1% on RAGTruth, both settings where the judge underestimates true factual accuracy.

ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation

Siyuan Zhang, Hanchen Wang, Dong Wen, Ying Zhang, Wenjie Zhang Retrieval-augmented generation (RAG) reduces hallucination but dense retrieval handles multi-hop question answering poorly, and graph-based RAG, which does follow multi-step relations, pays for it with semantic drift and slow global graph traversal at query time. ISO-RAG projects the knowledge graph into a hyperbolic Poincaré ball offline to precompute per-node isoperimetric profiles, then uses those profiles to prune spurious edges at retrieval time so Personalized PageRank diffusion runs over a strictly local subgraph and converges quickly. Across multi-hop question answering benchmarks this yields average absolute gains of 10.0% in retrieval recall and 4.3% in downstream exact match while removing the latency cost of global traversal.

The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space

Jacob Brinton, Jannik Brinkmann, Mark Crovella, Aaron Mueller The interlingua hypothesis proposed here says large language models translate by reading a source sentence into a language-agnostic latent feature space and then generating the target sentence out of that same space, rather than learning pairwise source-to-target mappings. Three lines of evidence are offered: BLEU variation across language pairs is largely predicted by per-language competence with no pair-specific interaction terms; many internal components are causally involved in both monolingual tasks and translation; and fine-tuning on monolingual data alone recovers a large share of the translation gains obtained by fine-tuning on aligned parallel documents. The authors argue this reframes how translation ability can be measured and improved in general-purpose models.

Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation

Zhuoheng Li, Ying Chen Retrieval-augmented generation draws on corpora that can contain outdated, contradictory, or unreliable documents, and human reliability labels are too expensive to collect at corpus scale. TrustPropRAG builds a graph of document relations and solves an optimization problem that combines pairwise relations with a small set of human feedback labels, propagating trust scores multiple hops so that a limited number of costly reliability judgments extends across the whole corpus. The resulting scores drive both document selection and trust-aware answer generation, improving retrieval quality and exact-match accuracy over baselines while staying robust when feedback is sparse or noisy.

EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection

Guanzhong Sun, Junyi Ma, Yuxuan Wu, Yanzi Miao A model whose choices vary with the input rather than tracking its own output prior is often taken as evidence of task competence. Using a forced-choice signalling task abstracted from the board game Deception: Murder in Hong Kong, where the fit-maximising, posterior-maximising, and uniform-random reference strategies are all computable in closed form, the authors test seven language models under three scoring rules and find every one of the 21 model-by-rule cells reliably item-sensitive — yet 8 of those cells are statistically indistinguishable from random choice and 5 score worse than random, with item-sensitivity and distance from random correlating at only r = 0.30. They call this consistency without alignment and argue it undermines any evaluation resting on item-sensitivity, permutation consistency, or self-consistency without an independent reference; a literal-similarity baseline with no pragmatic reasoning beats most of the tested models.

Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs

Seungwoo Jung, Dohyeok Kwon, Seungmin Cha, Junseok Lee, Yeonho Yoo, Chuck Yoo et al. Mixture-of-experts models are often shrunk by residual sparsification, which splits every expert projection into a shared base matrix plus a per-expert residual and then compresses each residual to minimize its own reconstruction error. The authors show that per-matrix error is the wrong objective, because an expert's output couples several projections and hidden representations, so small isolated errors compound into large output errors. PARSER retargets compression at expert output error using an output-importance measure of each weight's actual contribution, narrowing the accuracy gap to the uncompressed model by 1.41 times on Qwen and 1.44 times on DeepSeek at identical peak memory savings.

Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random

Cris Huynh See companion entry.

Predicting Program Exit Code with LLMs and Programming Language Semantics

Lara Marinov, Aditya Thimmaiah, Jayanth Srinivasa, Junyi Jessy Li, Milos Gligoric cross-listed Code-capable models may succeed at generation while lacking real command of programming-language semantics, and it is unclear whether they apply formal rules even when those rules are supplied in the prompt. The Program Executability Prediction task (PrEx) asks a model to judge whether a program is semantically valid given its syntax and operational semantics, and to name the violated rule when it is not; the accompanying dataset is built by systematically transforming valid programs into invalid ones across human-written, model-translated, and fuzzer-generated splits. Under two semantic formalisms and two semantic shifts, open-source coding models lean on pre-training priors instead of applying the given rules, performing especially badly when the semantics are modified and degrading further as programs grow more complex.

Enoki: Efficient Multi-Level Hallucination Detection

Elisei Rykov, Timur Ionov, Nikolay Ivanov, Maksim Savkin, Maksim Makarenko, Alexander Panchenko et al. Hallucination detectors generally work at one granularity — claim-level checks give interpretable factual units, span-level checks localize the offending text — and combining them costs extra decomposition, verification, and claim-to-span alignment. Enoki uses open information extraction to pull text-anchored relational facts, verifies each against evidence, and projects unsupported facts back onto spans, so one shared representation serves both levels without any separate alignment step. Extraction can run in language-model, encoder, or rule-based regimes behind a common interface to trade accuracy against cost; the system stays competitive with strong claim-level pipelines at lower resource use and leads on fine-grained span- and entity-level localization, alongside a released dual-granularity dataset called EnokiQA.

Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation Reranking

Guangyu Chen, Boxuan Lyu, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura Reranking methods such as Minimum Bayes Risk (MBR) decoding and Quality Estimation (QE) reranking improve neural machine translation output but require generating and scoring many candidates, and prior acceleration work targeted only the reranking step of MBR. Quit (Quantifying Uncertainty for Incremental Termination) treats candidate generation as a sequential decision under uncertainty, generating and reranking candidates incrementally and stopping once the top estimated quality score stabilizes. Across three translation models and 19 language pairs, it delivers end-to-end speedups of 1.47–2.66× for MBR and 3.43–4.12× for QE reranking while keeping quality inside prespecified equivalence margins.

Topological Steering

Beno\^it Gu\'erand, Tan Minh Nguyen Activation- and feature-space steering of LLM behavior operates on local geometry, which makes it sensitive to outliers, noise, and distribution shift. Topological Steering instead represents activation spaces through persistence diagrams borrowed from Topological Data Analysis (TDA), which capture global structure, and uses that representation to drive behavioral interventions. The authors report that the method consistently changes model behavior across multiple model families and sizes.

Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs

Jainil Dharmil Shah Edge deployment decisions are usually made on one metric at a time, so the authors combine accuracy, latency, memory, energy, and safety into a Holistic Sustainability Score organized around economic, environmental, and social pillars. Thirty measured configurations cover five natively small models at BF16 plus five larger models at BF16, INT8, NF4 4-bit, GPTQ 4-bit, and GGUF Q4, evaluated on five zero-shot benchmarks, GPU energy and throughput measurements, and attack success rate on harmful prompts. Qwen3-30B-A3B at GGUF Q4 ranks first overall at 93.38, ahead of Mistral-Small-24B at GGUF Q4, with Phi-4-mini the top-ranked small model, so the assumption that natively small models are always the more sustainable edge choice does not hold universally. The authors stress that quantization behaves as a systems-level choice rather than a smooth precision-efficiency trade-off, and that the score is relative to its comparison pool.

SCoNE: Selective Context-aware Neuron Editing for Robust Retrieval-Augmented Generation

Chaewon Kim, Seo Yeon Park Retrieval-Augmented Generation degrades when retrieved documents mix informative and irrelevant context, since the model gets distracted and hallucinates. SCoNE is a training-free model edit that locates feed-forward neurons scoring high on both attribution and cross-input variability — the ones treated as context-aware — and selectively strengthens them, using only a small set of mining samples, no fine-tuning, and no added inference cost. Across several knowledge-intensive question answering benchmarks and two LLM backbones it consistently outperforms competing noise-robustness baselines, with code released.

Value Over Language Model: Detecting Original Contribution in Writing

Vibhhu Sharma, Thorsten Joachims, Sarah Dean Detectors of machine-written text measure how much of a document's surface prose came from a model, not how much of its information content or ideas the human actually supplied. The VOLM (Value Over Language Model) framework never scores surface text at all: it extracts a document's content at increasing levels of granularity, has a language model reconstruct the document from each partial representation, and compares those reconstructions against ones generated from the task description alone, giving a contribution score relative to a replacement-level document. Across news articles, ICLR peer reviews, and argumentative essays it separates human-authored documents from matched model-generated ones while staying largely invariant to content-preserving transformations, including LLM rewriting and round-trip translation. Tightening the content extractor shrinks the residual gap between model-generated and humanized text, which the authors read as evidence that content must be disentangled from style.

Online Self-Weighted Fine-Tuning

Haiquan Wen, Yiwei He, Bei Peng, Guangliang Cheng Supervised fine-tuning (SFT) weights every expert demonstration equally regardless of whether the model already handles that query, while reinforcement learning adapts update strength but needs far more sampling and can destabilize on hard tasks. Online Self-Weighted Fine-Tuning (OSW-FT) keeps the gradient direction anchored to the expert trajectory but rescales the SFT loss per query by a success rate estimated from a handful of inference-only rollouts, an estimator the authors show is unbiased for the surrogate update at any finite rollout count and connect to SFT and RL through variance-reduction arguments. Across Qwen3 models from 0.6B to 4B on benchmarks including AIME, the method consistently beats SFT on small and medium models using only 2 online rollouts per query.

Can Large Language Models Forecast What Researchers Study Next?

Fenghai Li, Zihan Tang, Haofei Yu, Yining Zhao, Jiaxuan You Judging generated research ideas for novelty or feasibility at the moment they are produced says nothing about whether they anticipate what a field actually does next. IdeaForecastBench poses that forecasting task directly: given a community's literature up to a cutoff, a system ranks up to five ideas, which are scored against papers that appeared afterwards, over 624 rolling episodes across 52 topics with a fixed retrieve-then-judge protocol and two separately reported judges. Comparing five history-compression strategies over GPT-4.1, Qwen2.5-7B/14B, and Qwen3.5-9B plus a learned Mode-Decomposition Forecaster, summarizing the history beats feeding it directly on Hit@5 and Precision@5 for all four backbones, with Qwen2.5 scoring above GPT-4.1 by producing broader forecasts; threshold and judge diagnostics temper how much realization should be read as precise anticipation.

How Do Language Models Choose Between Context and Memory?

Benjamin Shih, John Winnicki, Arianna Cao When retrieved context contradicts what a model learned in its weights, activation directions can steer which source wins — but steering along a direction does not show the unedited model uses it, or that it transfers. The authors estimate authority directions from agreement prompts where context and parametric knowledge concur, then swap naturally occurring coordinates along those directions between matched prompts instructing the model to favour one source or the other. Across Qwen, Llama, and OLMo, the swap reproduces 30-68% of the authority-induced shift in source choice while matched controls reproduce almost none; directions learned on one task close only 9% of the authority gap on another versus 57% for the locally learned direction, suggesting the computation is task-dependent rather than a reusable global feature.

Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning

Jinhu Qi, Minda Hu, Wentao Zhang, Weiqiang Jin, Yanyu Chen, Junli Wang et al. On in-context learning tasks where a long novel context defines rules, knowledge, and output schema and grading checks every detail, strong open-weights models pass only 12-16% of tasks because one missed rule sinks the whole response. The authors argue the read-and-reason paradigm is structurally to blame — extraction, planning, generation, and self-verification all crammed into one forward pass — and propose the Context Compilation Architecture (CCA), which compiles prose context once into a typed intermediate representation with fixed slots for must-do, must-not, and conditional rules, output spec, available tools, and data profile, then runs executable verifiers and a violation-gated correction loop. Across 1,899 tasks in CL-bench and four open base models it beats vanilla prompting and both long-context baselines (ReadAgent-P, Ctx2Skill) on every model, lifting Kimi K2.5 from 15.4% to 21.4% with gains concentrated on rule-dense categories.

Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation

Wentao Ye, Zhanming Shen, Zhiqing Xiao, Yao Ding, Haobo Wang, Gang Chen Parameter-efficient fine-tuning is normally debated in terms of how many parameters to train, but under a severe budget where trainable factors cannot repair a bad subspace, where those coefficients sit matters just as much. The authors study this with frozen-core adaptation — a calibration pass fixes left and right bases per weight matrix and only an r x r core is trained — and propose FCCA, which estimates the signed input-error cross-covariance, whitens it with diagonal Fisher moments, truncates in that local metric, maps back, and applies thin QR for stable core coordinates. Comparing eight basis constructors on 11 tasks, four model settings, and three seeds, FCCA leads at all three Qwen scales (83.0 macro-average on Qwen2.5-3B, 2.3 points above the next matched-budget method) and lands within 0.32 points of LoRA while training 36.9K parameters instead of roughly 7.4 million; ablations attribute 2.7-17.2 points to whitening and find QR necessary for stable optimization.

Instella-MoE Technical Report

Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra, Yonatan Dukler, Gowtham Ramesh, Jialian Wu et al. Instella-MoE is a fully open Mixture-of-Experts (MoE) language model with 16 billion total and 2.8 billion active parameters per token, trained from scratch entirely on AMD Instinct MI300X and MI325X GPUs. The architecture combines sparse expert routing with Gated Multi-head Latent Attention and FarSkip-Collective connectivity, and the training pipeline runs from pre-training through mid-training, long-context extension, supervised fine-tuning with feedback-driven data curation, direct preference optimization, and reinforcement learning with multi-teacher on-policy distillation. It averages 76.7 across standard pre-training benchmarks, ahead of other fully open models such as OLMo-3-7B and OLMoE-1B-7B while staying competitive with open-weight models like Moonlight-16B-A3B and Qwen3.5-4B, and the post-trained Think checkpoint averages 73.2 across instruction-following, reasoning, math, coding, and chat. Weights, training configurations, data mixtures, and training code are all released.

SFAD: Speculative Factuality-Aware Decoding

Guanqiao Chen, Di Wang, Lijie Hu Keeping generated text faithful to supplied context is usually bought with extra compute: contrastive decoding needs two forward passes per step, and post-training alignment needs substantial reinforcement learning. SFAD (Speculative Factuality-Aware Decoding) folds faithfulness into speculative decoding instead, training a context-faithful draft model by Direct Preference Optimization on ConFide, a preference dataset built from fine-grained atomic perturbations. At inference an Epistemic Friction score quantifies distributional tension between draft and target weighted by draft certainty, and when it crosses a threshold an asymmetric residual-based logit injection steers the target distribution; otherwise normal speculation continues. The reported result is improved faithfulness alongside a 2.48x speedup rather than the slowdown contrastive methods incur.

Towards a Reliable and Practical Eval Pipeline

Emma Thuong Nguyen, Abhishek Ghose Teams shipping LLM-based software increasingly gate releases on "evals", but published work tends to address single aspects of eval reliability rather than what a production pipeline actually needs. The proposed end-to-end pipeline pairs automated creation of eval checklists with a learned aggregation step over the checklist responses, improving both agreement across LLM judges and accuracy against human judgments. It also surfaces self-consistency, explanations, and prediction uncertainty, with empirical results supporting the design.

Replacing Training with Memory: Listwise Selection for Text-to-SQL

Yeonseok Jeong, Soyoung Yoon, Seongjun Lee, Seung-won Hwang cross-listed Text-to-SQL pipelines that generate many candidate queries and pick one usually rely on a listwise selector that must be fine-tuned, which is expensive. MaP-SQL replaces both fine-tuning objectives with inference-time machinery: reusable structured memories distilled from training data encode how natural language maps to schema elements, SQL operations, and expected outputs, serving as explicit criteria for comparing candidates, while rankings are aggregated across multiple input permutations to cancel positional bias, with execution results and pointwise scoring keeping the comparison count down. On BIRD-dev with the same candidate sets, it beats the prior selector-based state of the art R^3-SQL by 2.02 execution accuracy points while using 2.92x fewer tokens.

CacheBridge: Efficient Cross-Model KV Cache Transfer

Xingyu Qu, Siyuan Lu, Zhiyu Chen, Sheng Wang, Tao Lin When several large language models share context in a multi-model system, the receiving model normally has to re-prefill the shared prefix because key-value (KV) caches are model-specific. A recent training-free approach, Full-Head Mapping, fits a closed-form affine mapper between source and target caches, but maps every target head from every source head, making it fragile to architectural differences and expensive to store and apply. CacheBridge restricts each target head to a single matched source head, weights reconstruction errors by causal attention sensitivity, and builds the mapper with a fused GPU kernel that avoids materializing full observation tensors. It recovers two Ministral 3 transfer directions where Full-Head Mapping loses substantial accuracy, keeps 99.83% mean target retention on Qwen3, and on the Qwen3 14B-to-32B transfer cuts mapper storage by 8x and 500-sequence construction time from 92.63 to 8.63 seconds while matching the baseline with a tenth of the calibration data.

Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO

Prakhar Gupta, Vaibhav Gupta Language models often ignore prompt evidence that conflicts with memorized knowledge, and post-training can improve this, but it is unclear whether the gains build new machinery or amplify what the base model already has. The authors compare nine post-training arms spanning GRPO, supervised fine-tuning (SFT), and DPO from a single starting checkpoint, extended across scales and families, and estimate a grounding direction from that checkpoint before any training. Five GRPO variants yield small grounding gains even as the rewarded metric improves, conflict-SFT helps moderately, and DPO pushes grounding near ceiling on its matched distribution, yet both SFT and DPO use largely the same causal attention-head set as the starting model. Subtracting the pre-existing direction suppresses both gains, adding it to the base model recovers 35% of DPO's gain at a dose passing all side-effect checks, and a supervised warm start leaves GRPO with essentially no further grounding to add, indicating the gains largely depend on machinery already present.

PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition

Ziyan Gan, Fangxin Liu, Chenyang Guan, Junjie Wang, Ning Yang, Haomin Li et al. Mixture-of-Experts (MoE) inference frameworks treat each expert as an atomic execution unit, which fixes the optimization boundary too early and ignores computational redundancy inside experts. PCoMoE reframes MoE inference as fine-grained path composition, introducing a path-level formulation of expert computation, a compatibility-aware layer-wise pruning strategy that suppresses low-value path combinations, and a hardware-friendly execution engine that exploits reusable sub-expert structures under bounded overhead. The authors report up to a 1.31x end-to-end inference speedup alongside a 10% accuracy improvement, with code released.

Lagged Coupling: Internal Representations Become Readable Before They Become Causal

Xining Xun Probe accuracy is often taken as evidence that a model's internal representation can be used for steering, and the authors test that assumption across the full Pythia suite (160M to 12B, eight checkpoints, four task families) plus a pre-registered OLMo-2 replication. A linear probe reads the target variable from the residual stream at AUROC of at least 0.990 from step 1,000 at every scale, yet steering along that same direction is null-equivalent in 43 of 48 model-checkpoint cells, and the lag does not shrink with scale. They name this structure lagged coupling and decompose it into three tracks (internal readability, behavioral readability, and causal efficacy), finding that representation headroom grows up to 57x with training while causal write-in stays under 0.11% of it. Two pre-registered single-onset hypotheses resolve as indeterminate, and the authors caution against inferring steerability from probe accuracy.

OUTLETS: Output-Length Prediction from Speculative Decoding Backbones

Weihuang Wen, Yingying Liu, Yichuan Liu, Wenqi Zeng, Li Zhou, Chumin Sun et al. Heavy-tailed output lengths in Large Language Model (LLM) serving complicate resource provisioning and scheduling, and existing length predictors either add latency through external proxy models or rely on shallow probes of current model state. OUTLETS observes that the latent representations produced by the draft decoder in speculative decoding frameworks such as EAGLE-3 already encode signals predictive of generation length, and attaches a lightweight regression head to that backbone to turn it into a trajectory-aware length predictor at almost no added cost when draft states are computed anyway. It achieves lower mean absolute error than the evaluated methods, and under saturated disaggregated serving its predictions let standard scheduling policies prioritize short requests and balance load across decoding instances, cutting short-request P99 latency by 34.8%.

Post-hoc Alignment of LLM-judges to Human Judgment Distribution

Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani LLM-as-a-judge (LLMaJ) evaluations are typically scored against aggregated ground-truth labels, discarding the Human Label Variation (HLV) that a distribution of annotator judgments carries. Across five datasets the authors find LLMs approach human-level accuracy at predicting a single hard label but perform poorly at predicting the soft-label Human Judgment Distribution (HJD). NAPHA (eNtropy-Aware Post-Hoc Alignment) is a lightweight fix that assigns each instance to a discrete entropy class and routes it to a specialized trained alignment model that maps the LLM's distribution onto the human one; it consistently improves soft-label prediction across base models and datasets, with the largest gains on high-entropy instances, and oracle experiments show better entropy-class prediction would raise its effectiveness further.

Beyond Magnitude: Contrastive Routing for Modular Mixture-of-Experts

Nikolaos Xiros, Dimitrios Damianos, Maria-Eleni Zoumpoulidi, Leon Voukoutis, Vassilis Katsouros, Georgios Paraskevopoulos Routing in Mixture-of-Experts (MoE) layers operates on token representations dominated by structure shared across all tokens, which limits how much experts can specialize. The Contrastive Routing Mechanism (CoRM) contrasts each token against an Exponential Moving Average of the layer's hidden states and scores each expert by the gap between its affinity for the token and its affinity for that shared reference state, through a distinct per-expert projection; the authors show this concentrates the routing signal into a low-dimensional, highly separable subspace with boundaries that align more closely with linguistic structure than standard Top-k routing. On nine zero-shot reasoning benchmarks CoRM improves average accuracy by +0.67 to +1.69 points (Top-1) and +1.38 to +1.77 points (Top-2) over Top-k MoE baselines, at a cost of 2.9% more parameters and 2.6% more FLOPs per token.

Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation

Will Badr cross-listed When a hint turns a failing generated program into a passing one, it is unclear whether the hint supplied missing information or merely steered the model toward a solution it could already reach. The authors test this with executable evaluation on HumanEval+ and MBPP+ using Qwen2.5-3B-Instruct and Phi-3.5-mini, comparing adaptive relevant hints, an unrelated hint, and plain repeated sampling without hints, and they add mechanistic probes of a hint-related activation direction. Relevant hints rescue 36 of 79 selected Qwen failures, but eight unhinted samples recover 31 of those 36, and the same pattern holds for Phi, where unhinted sampling recovers 36 of 42 relevant-hint rescues. Persistently adding the shared hint direction yields 14 rescues and 18 regressions with no detectable net gain, so the evidence does not establish task-general capability transfer, though the authors note that differing attempt budgets across conditions prevent isolating a purely semantic effect.

Does task decomposition improve automatic NLG evaluation?

Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani The LLM-as-a-judge (LLMaJ) framework offers cheap, reproducible, reference-free evaluation of Natural Language Generation (NLG), and prior work has tried to improve it by decomposing evaluation into simpler sub-tasks. This study systematically compares LLMaJ methods with and without decomposition across multiple NLG datasets against a fair baseline that does not decompose. There is no evidence that task decomposition improves performance; previously reported gains stem from using human labels as training data rather than from decomposition itself. When human labels are available, LLMaJ without decomposition performs comparably to human annotators.

Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training

Guangqi Li, Yongxin Li Large language models show a modular internal organization resembling functional networks in the human brain, but prior work has only characterized finished models rather than how that organization forms. The authors train Pythia-410M from scratch in two trajectories (bf16 and fp32) and run attribution patching at every step, alongside probes of gradient norms, effective updates, weight norms, and first-order loss decomposition across 14 tasks in four cognitive domains. The modular map is pre-carved: before any learning, the dominant task pair already overlaps at about 3.6 times the task-independent attribution baseline, and the partition then locks in through two sharp jumps whose amplitudes do not track the learning-rate schedule, accompanied by gradient-level relative deprivation in which winning tasks receive 2.25 to 2.73 times the loser's gradient supply without this propagating to updates or weights. Deviation from the baseline substrate appears only in the domain being learned, and the authors pre-register a scale-threshold hypothesis for ongoing 2.8B experiments.

LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs

Muhammed Saeed, Simon Razniewski Flagship language models score above 90% on benchmarks like MMLU, yet fixed question sets test only what experimenters thought to ask. LLMPEDIA recursively materializes about 1.3M encyclopedia articles from the parametric memory of GPT-5-mini, DeepSeek-V3.2, and Llama-3.3-70B without retrieval, then audits a stratified sample of atomic claims against Wikipedia and a curated web stack, labeling each claim supported, refuted, or insufficient. On a uniform random sample only 68.4% of claims are true, more than 21 percentage points below MMLU scores, and 30.5% are insufficient, meaning neither benchmarks nor the world's largest encyclopedia can adjudicate them, whether long-tail knowledge or plausible hallucination. The result is a live, open encyclopedia offering link-traversal exploration, claim-level factuality, cross-model and political-persona comparison, and guided topic drill-down, with every page, claim, and verdict at a stable URL.

FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue

Hangyeul Lee, Juyoung Oh, Jaeyong Ko, Sunmin Kim, Jaeik Park, Hyunkyu Kim et al. Banking assistants that interact with the same customer repeatedly must keep complete, current, and traceable records as life changes surface incidentally in routine requests, but existing benchmarks test question answering or bounded recall rather than exhaustive longitudinal reconstruction. FinLifeBench contains 6,000 eight-turn Korean banking sessions from 20 synthetic customer trajectories and asks models to reconstruct every life-event instance with its first-establishing session and to rebuild a complete 34-path financial state at consecutive checkpoints, with deterministic gold labels for 24 event types. Across eleven LLMs given the full context, event-anchor recall falls from 0.591 at 15 sessions to 0.445 at 300, with errors driven mainly by omitted events, while financial-state reconstruction frequently treats superseded information as current and the best checkpoint accuracy reaches only 0.470. Performance on the two tasks is only weakly correlated, indicating models can localize evidence for events they recover yet still fail to maintain temporally valid records.

Prompt-Robust Language Models: Which Training Strategies Work?

Frederic Sadrieh, Michal \v{S}tef\'anik Large language models remain highly sensitive to how a prompt is phrased, and prior work has tackled this either through refined data construction or dedicated robustness objectives. The authors reproduce and compare these strategies under controlled conditions and find that robustness fine-tuning beats standard fine-tuning and in-context learning, yet the gap between the best and worst prompt template remains 40 to 57% of performance. Recent objectives such as CoIN for contrastive alignment and PPCL for consistency regularization often fail to beat the simplest data strategy of training on one template per batch. Diagnostics show the auxiliary objectives only move the quantity they penalize without generalizing, and that per-template gradients conflict in sign on 57 to 64% of parameters, so mixed-template batches force the optimizer to reconcile competing updates instead of finding a shared prompt-agnostic one.

Post-Training Science for Supervised Fine-Tuning

Charles O'Neill, Mudith Jayasekara, Harry Partridge Every supervised fine-tuning (SFT) run rediscovers the same choices from scratch: learning rate, batch size, LoRA versus full fine-tuning, epoch count, optimiser, and data. The authors run a controlled sweep that varies one lever at a time across dense and mixture-of-experts models from the Qwen3 and Llama families, on four real-world customer SFT datasets whose training data was iteratively refined to pass a customer-built evaluation, for both LoRA and full fine-tuning. The study asks how optimal learning rate and batch size shift with scale, family, and data and whether one selection rule transfers; what LoRA trades against full fine-tuning and how rank and alpha bound what an adapter can learn; whether validation loss or loss-landscape flatness faithfully ranks downstream quality; how gains scale with model size and data on a ladder reaching 235B parameters; how many epochs can run before general instruction-following erodes; and whether a geometry-aware optimiser beats AdamW. Each recommendation is paired with a measure of its uncertainty.

mzCache: On-Device LLM Memory Management under Multitasking

Hongseung Yu, Minsung Kim, Jongseok Park, Kyunghan Lee cross-listed Phone users switch apps constantly, so the operating system evicts model weights and key-value cache under memory pressure, forcing an on-device language model to reload from slow storage or recompute its whole cache when the next request arrives. mzCache manages memory for this multitasking regime by splitting model memory into fine-grained shared buffers that support partial eviction and restoration, and by using the unified memory of mobile systems-on-chip to keep GPU inference running while the CPU restores in parallel, with hybrid swap and backward-out eviction policies. Implemented on llama.cpp and deployed as an Android application, it cuts time-to-first-token by 2.1 to 5.5 times relative to storage-backed partial offload.

Probing Factual Knowledge Transfer with Training Data Interventions

Romina Oji, Marc Braun, Marcel Bollmann, Marco Kuhlmann, Jenny Kunz Whether multilingual models genuinely carry facts across languages or mostly recall facts seen in the target language is hard to test observationally, so the authors intervene on the training data: starting from an English-pretrained model, they continue pretraining on Persian text with specific facts systematically removed at several levels of granularity. Their SIFT resource contains 500 triples across 20 relations, split by whether the subject is globally prominent or Persian-specific, with natively written Persian cloze templates. Under the strictest removal condition a large majority of English-acquired facts fail to transfer into Persian, sentence-level co-occurrence filtering leaves fact signal intact, and randomly chosen negative candidates inflate apparent transfer by rewarding shallow associative heuristics. Facts about Persian-related entities, far rarer in the English corpus, barely transfer at all.

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Shaowen Wang, Ge Zhang, Kairong Luo, Yuhao Wu, Shaofan Liu, Jiaheng Liu et al. Looped transformers add effective depth by iterating a shared block, but comparing at fixed model size hands the looped variant extra floating-point operations, conflating architecture with compute. SMELT loops the middle half of layers twice in a Mixture-of-Experts transformer while matching per-token FLOPs, total non-embedding parameters, and key-value cache against an unlooped baseline, scaled across four sizes up to 54B non-embedding parameters with a separate Chinchilla-style scaling law fit per architecture. Loss falls faster with compute, saving 6.8 to 18.0 percent of training FLOPs on the compute-optimal frontier, with downstream gains exceeding what validation loss predicts, largest on code, and growing with sequence length and in-context example count. Mechanistic analysis credits the second visit with shrinking the attention sink and redirecting attention toward content-relevant tokens.

Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades

Dushyant Rajput Inference cascades route most queries to a cheap model and escalate a hard tail to a frontier model acting as verifier, and a tempting extension finetunes the cheap student on the verifier's rejections so escalation and cost fall each round. Measuring the loop on real models, the authors find the verifier's blind spot, the share of student errors it wrongly accepts, grows with student capability (0.12 at 0.5B parameters to 0.55 at 32B) and shrinks with verifier capability, so it is worst in the cheap-student, cheap-verifier regime cascades exist to create; a frontier verifier nearly closes it but then escalates 46 percent of hard MATH queries against a 39 percent true error rate. Corrective finetuning on the rejected tail degrades and eventually collapses the small student across every teacher tried, cross-family and same-family alike. Throughout, every metric computed through the verifier reads a flat 3 percent error while true delivered error swings as high as 32 percent, a blindness the authors formalize as a two-population conservation law and validate synthetically.

CHARM: Character Hallucination for Multicultural Role Play Benchmark

Sunkyung Han, Nahyeon Park, Gaeun Seo, Seunghyun Yoon, JinYeong Bak Role-playing language models are supposed to hold a character's voice while respecting that character's knowledge limits, but existing hallucination evaluations do not separate failing to notice a boundary from crossing it after noticing. CHARM covers 40 real and fictional characters from five cultural-linguistic regions, validated by native reviewers, probing temporal and cross-universe boundaries with abstention-enabled multiple-choice questions and a two-stage score that splits boundary awareness from boundary compliance. Across six models, hallucination is driven predominantly by compliance failures: the model states that a question lies outside the character's knowledge and then answers it factually anyway. Re-posing the same questions to the character confirms many cases are parametric overrides where the fact is stored but not suppressed, and failure rates vary systematically by cultural region.

Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs

Mikhail Sonkin, Tanja Baeumel, Daniil Gurgurov, Josef van Genabith, Simon Ostermann Multilingual models translate well, yet how they transform a representation from one language into another is only partly understood; prior work splits the process into language-independent conceptual content followed by production into language-specific form. Using controlled multilingual datasets that isolate word-order differences, plus causal interventions and probing, the authors show the production stage splits further, with models committing to target-side word order before realizing the target language's surface form. They also identify individual attention heads that respond selectively to syntactic transformations while remaining largely invariant to language identity, placing syntactic commitment as its own stage in the translation pipeline.

Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA

Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto Linear classifiers trained on a model's hidden states can flag factual errors in a single forward pass, implying true and false statements separate along a stable truth direction, but published results disagree on whether that direction survives input shifts because cross-dataset transfer experiments change several variables at once. The authors isolate writing style, medical specialty, and source corpus by rewriting 500 MedQA items into textbook, patient, clinical-note, and colloquial registers, annotating each with a specialty, and grouping them with MedMCQA and MMLU-medical. Probing four open-weight models of 2 to 8B parameters, style costs about 0.10 AUROC and specialty about 0.03, while corpus shift costs up to 0.21 AUROC, roughly twice the register gap. The register result replicates with a second generator and with human-written patient questions, so question format does not explain the corpus break, suggesting the probe signal is partly bound to dataset structure rather than medical knowledge.

How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation

Elitsa Yotkova, Violeta Kastreva, Petar Velkov, Hristo Boyanov, Dimitar Dimitrov, Ivan Koychev et al. Scoring free-form question answering is a bottleneck because a correct answer takes many surface forms and a wrong one can fail by incompleteness, contradiction, overgeneration, or endorsing a false premise, distinctions that judge-based and similarity-based metrics collapse. The authors define an eight-class ordered taxonomy of semantic correctness that keeps verbose-but-correct answers apart from answers carrying hallucinated content, and release CAP-Correctness with 8.8k examples across widely used question-answering datasets plus CAP-Statements with 11k question-answer-to-statement conversions for natural language inference (NLI) training. Their reference-based metric CAP (Context-Aware Precision) scores question-conditioned statements with bidirectional NLI and outperforms established baselines under a monotonicity protocol that tests whether a metric respects the taxonomy's intended ordering.

Behaviorally Effective LoRA Writes Are Sparse and Structured

Haruto Sato, Yuki Tanaka, Ren Nakamura, Aoi Kobayashi, Mei Ito Low-rank adaptation fixes the rank of an update but says nothing about which parts of a trained adapter actually carry behavior. The authors warm up an unconstrained adapter, convert its learned write columns into a frozen module-wise orthonormal basis, and continue training inside that constrained parameterization; across 14 exact switches held-out accuracy is unchanged at conversion and reconstructed write matrices differ by at most 0.25 percent relative Frobenius error, while continuing the same checkpoint under different write subspaces leads to different outcomes, marking write geometry as a causal state variable. A no-retraining projection test shows the useful signal stays inside the learned write space and vanishes in random or frozen-activation principal-component controls. On GSM8K, MathQA, and AQuA, per-module top-k continuation peaks at k of 2 or 4 in all twelve seed-level cases, learned top-16 and top-32 subsets beat matched random subsets, and single-direction ablations isolate a few late query, output, and down projection components with outsized behavioral impact.

When Tokenization is Secretly Output Supervision

Tanja Baeumel, Josef van Genabith, Simon Ostermann Tokenization is normally treated as an input preprocessing choice, but in autoregressive models the tokenizer also determines what must be resolved in a single forward pass and therefore what supervision signal the model receives. A controlled numeric-reasoning experiment that decouples input from output tokenization finds that differences in task performance, training dynamics, and model internals are induced by output tokenization and are largely invariant to input tokenization, implying that models with different tokenizers were effectively trained on different tasks rather than differing only in ability. A survey of 120 recent computational-linguistics papers on numeric reasoning finds only about 10% report the numeric tokenization of the models they evaluate, while 69% compare across tokenization regimes without reporting it.

Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

Zhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat, Eytan Bakshy, Jihao Andreas Lin Optimal hyperparameter scaling laws let practitioners predict good configurations at production scale, but fitting them conventionally means exhaustive grid searches over thousands of training runs. Power-Law Entropy Search (PLES) is a cost-aware acquisition function for multi-fidelity Bayesian optimization that selects, at each iteration, the configuration that most reduces uncertainty in the scaling-law estimate per unit of compute, targeting the law itself rather than a single objective and naturally favoring cheap small-scale experiments. Across synthetic benchmarks, surrogates fitted to real large language model training data, and actual pretraining runs, it converges to accurate scaling laws using less than one-tenth of the compute required by grid search and other baselines.

Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation

Yixuan Liu, Lin Chen, Zhuoqi Liu, Jianglin Lu, Dakota Murray cross-listed Citations carry rhetorical intent — supporting, contrasting, or merely mentioning prior work — and whether models reproduce that intent when assisting scientific writing is untested. A masked-citation task has six popular large language models generate replacement citation sentences, producing a counterfactual corpus directly comparable to human writing over 1,746 top natural language processing conference papers, 63k+ contexts, and 132k+ citations, with an LLM judge classifying intent and a 20-million-edge coauthorship network measuring social distance to cited authors. Models cite significantly less critically than humans, over-cite popular and older papers (a tendency strongest where humans would contrast against recent, niche work), and draw on more socially distant authors than the close collaborators humans favor for supporting citations.

LatentPress: Context Compression Beyond Text and Vision

Zhengze Zhou, Hejian Sang Compressed context is normally carried as readable text or rendered images that must be decoded, even when the consumer is a language model. LatentPress writes conversation histories and long documents into continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, training only a 4.2M-26.2M-parameter adapter (roughly 0.1% of the decoder) to compress 4-16x with no text reconstruction at inference. On LongMemEval it reaches 0.504 accuracy at 7.70x compression, above the 0.490 obtained from uncompressed evidence, and far above text summaries (0.184) or OCR-based compression (0.426 to 0.312); writing costs 43ms per conversation, about an order of magnitude faster than summarization or OCR, and reading is 5-9x faster than raw context. Transfer holds zero-shot from UltraChat to LongMemEval and onward to unseen LongBench document domains, though 16x compression still trails raw context.

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia et al. Logit-based knowledge distillation trains small language models from stronger teachers, but its benefits turn out to depend on the training stage. Forward Kullback-Leibler distillation improves both reasoning and factual recall during pre-training, yet during mid-training — the intermediate self-supervised phase on curated corpora — it keeps delivering reasoning gains while slowing factual recall, which the authors trace to teachers being more confident on procedural than knowledge-intensive data while students acquire low-entropy facts early. Switch Distillation uses teacher predictive entropy as a routing signal, distilling only where the teacher is confident and falling back to cross-entropy elsewhere; relative to standard next-token prediction it achieves 1.61-1.71x the reasoning performance while preserving 96.7-96.8% of factual recall, and the advantage survives post-training.

Can LLMs Discover Scientific Laws in Real and Parallel Worlds?

Yiming Huang, Ziche Liu, Zhuohang Wu, Yiqian Wang, Junxia Cui, Xinkai Zou et al. Whether large language models can genuinely discover scientific laws is hard to judge because existing evaluations either use simplified synthetic setups or reuse published targets the models may already have memorized. SciLaws-Bench draws 118 problems from 381 papers, covering 291 candidate laws and roughly 8 million real data points across six disciplines, and poses each in two settings: SciLaws-Real, where models propose laws from fixed real observations and are judged on held-out predictive fit and literature-derived scientific validity, and SciLaws-Parallel, where models actively query residual-calibrated simulated worlds to recover a newly synthesized hidden law. Predictive fit can diverge from scientific validity, memorization determines whether models reproduce or move beyond published formulas, and a best-of-N study reveals a selection bottleneck.

Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories

Nabira Rashid, Manolis Kellis Embedding retrieval is tested where surface form and meaning are deliberately pulled apart, retrieving items that share underlying structure but not wording, on competition mathematics (MathNet-Retrieve, 500 queries over a 117,088-item corpus) and ALFWorld-derived embodied-agent trajectories (118 queries, 336 trajectories). In mathematics, strict Hit@1 at the heaviest disguise tier is 0.0% for both production embedders, even though the correct item almost always sits in the top 10 and the winner is more lexically similar to the query in 95.2 to 99.8% of misses; in trajectories the same models fall to chance or below once the gold item must differ in object and receptacle. A lexical reranker hurts in mathematics but helps in trajectories, which the authors use as a diagnostic for whether a benchmark's surface variation is adversarial or incidental, while an LLM reranker recovers 5 to 63% of the gap in mathematics and 43 to 76% in trajectories, with part of the mathematics gain traced to memorization of well-known competitions. A paired downstream experiment found oracle retrieval indistinguishable from adversarially bad retrieval because the solver's zero-shot accuracy was largely a truncation proxy, leaving no headroom for retrieval to matter.

From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification

Manish Gupta, Chaitanya Giri, Jayasimha Talur Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, and the usual fix of retrieving the top-K candidate labels by embedding similarity narrows the choice without helping the model tell near-duplicates apart. The proposed framework identifies which label pairs the model confuses, expands the candidate set to include those confusable labels, and generates targeted rules that distinguish similar candidates, all without fine-tuning. On WOS, Flipkart, and LEDGAR, Macro F1 improves by up to 10.0 percentage points over retrieval baselines, and the generated rules transfer to smaller 2B to 20B models, which gain up to 11.5 points.

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

Olga Tsymboi, Dmitrii Stoianov, Ramil Latypov, Danil Taranets, Daniil Dryabin, Mikhail Gashkov et al. Data-residency rules push enterprises to self-host large language models, and adopting newer models without retiring their predecessors fragments a finite GPU pool across a growing serving fleet. The authors consolidate traffic from over 200 internal applications onto a single model by closing quality gaps found through production error analysis along instruction following, function-calling, and the internal task distribution, tracked with offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimizing all objectives jointly, which caused cross-domain reward interference, they train a separate GRPO expert per axis and merge them with two-stage SLERP, with each expert's reward exposing a distinct failure mode: semantic collapse, over-calling, and verbosity hacking. In non-reasoning mode the model beats a roughly 7x larger baseline on an in-house Arena (69.6 vs 65.8), instruction following (0.85 vs 0.83), and function-calling (0.79 vs 0.77), and it now absorbs 50% of platform traffic, 116 million requests per month, at a fraction of the serving cost.

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

Jingtan Wang, Arun Verma, Xiaoqiang Lin, Zhengyuan Liu, Nancy F. Chen, Daniela Rus et al. How to split a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) in LLM post-training is unresolved, with prior work offering only broad trends and no evidence on whether the best ratio transfers across model sizes. Instead of hunting for a single optimal ratio, the authors characterize the near-optimal region, the set of allocations within a given tolerance of peak performance. That region is wide even at 2 to 10% tolerances, widens with model scale, and transfers reliably from small proxy models to large targets, so cheap proxy experiments suffice to pick an allocation without exhaustive large-scale search. The pattern holds across tasks, model families, and both preference-based off-policy and reward-supervised on-policy RL methods, and the authors show how asymmetric annotation costs for SFT versus RL data shift the region.

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

Jundong Hu, Shekar Ramachandran Post-training quantization (PTQ) cuts LLM serving cost, but its accuracy damage is uneven and typically tuned per model. Using causal mixed-precision intervention as ground truth, raising each layer to 8-bit in turn and measuring recovered accuracy across 9 open-weight models from 4 architecture families, the authors test whether damage lives in task circuits, where the model computes, or in weight statistics, and find that none of these predicts which layers benefit from restored precision. Recovery is instead diffuse: for 8 of 9 models, recovering 75% of the gap takes roughly half the layers, with Qwen3-8B the lone sharply concentrated exception. At a matched precision budget, spending it globally on finer quantization granularity beats locally repairing the most recoverable layers by 21 to 52 points for every group-128-compatible model, and since 8-bit is near-lossless across RTN, GPTQ, and AWQ, the authors conclude that cheap correlates of quantization damage must be checked by causal intervention before guiding precision allocation.

Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation

Kefeng Duan, Dewu Zheng, Yanlin Wang, Terry Yue Zhuo, Mingwei Liu, Jianxing Yu et al. cross-listed Repository-level code generation must produce code consistent with a codebase too large to fit in a model's context, so most systems rely on retrieval-augmented generation (RAG) that supplies repository context as task-level support. ACToR instead identifies critical tokens, the few decisive positions during autoregressive decoding where an error sends the rest of the output down a wrong semantic path, and triggers targeted retrieval on demand at exactly those positions, aided by a position-aware weighting scheme that makes dense retrievers prioritize context most informative for generation. On RepoExec and CoderEval it consistently beats state-of-the-art methods, with relative improvements of 8.4% and 15.4% respectively, and an accompanying analysis quantifies how heavily major generation failures concentrate at these critical tokens.

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

Himil Vasava, Ming Jiang Large language model judges are widely deployed to score natural language generation (NLG) quality and to supply automated training signals, yet the internal procedure by which they assign a rating is poorly understood. The study probes this mechanistically with an eight-attack perturbation taxonomy over the Readability and Adequacy dimensions, a pipeline producing paired clean and corrupted summaries with controlled error intensity and token-level modification maps, and a battery of causal tracing, logit-lens projection, and attention-head knockout applied to Themis (Llama-3-8B) and Prometheus (Mistral-7B). Both evaluators implement a two-stage pipeline: below layer 15, attention performs local error comparison and routes the result to the final input position, while above it the MLP cascade integrates the signal and writes the rating, with the decision crystallizing sharply at layer 26 on Themis and layer 25 on Prometheus. A base Llama-3-8B control reproduces the routing and crystallization but not the stage separation, isolating two effects that fine-tuning specifically installs, suppression of early MLP contributions at the last position and a two-layer earlier crystallization, which indicates fine-tuning sculpts an existing substrate rather than building the pipeline from scratch.
19 more specialized papers

Applications 98

Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models

Parviz Ghafariasl, Weimin Fu, Xiaolong Guo, Shing I. Chang Scams aimed at older adults unfold over several conversational turns, escalating from casual contact through trust building and urgency to a request for money or credentials, so risk has to be re-estimated as the dialogue grows. The proposed framework accumulates turns incrementally and re-scores risk at each step, trained on a purpose-built multi-turn dataset of investment, charity, and tech-support scams annotated at every cumulative stage with a risk level, continuous score, rationale, and safety recommendation. Four compact models — Phi-4, LLaMA-3.2, DeepSeek-R1, and Qwen3 — were fine-tuned under a shared recipe, with Phi-4 and LLaMA-3.2 giving the strongest turn-aware risk estimates relative to their parameter count, supporting on-device deployment where conversations never leave the phone.

Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation

Ante Kapetanovic, Tomislav Duricic, Andro Mercep, Emanuel Lacic LLM rerankers for conversational recommendation are compared against collaborative-filtering and sequential baselines inside a shared retrieve-then-rerank pipeline on the ReDial movie benchmark, while varying candidate-pool size, first-stage retriever, and decoding temperature. With a shared semantic top-250 pool and strict candidate-aware scoring the best proprietary reranker reaches NDCG@10 of 0.1497 against 0.0939 for the strongest non-LLM baseline, yet the same model scores 0.2925 under unconstrained zero-shot generation, and no open-weight model beats a tuned shallow autoencoder. Switching from semantic to collaborative-filtering candidates raises NDCG@10 by more than 50% for the strongest rerankers, and higher temperature mainly changes list stability rather than mean accuracy for strong models, leading the authors to argue candidate generation, pool size, scoring policy, and decoding configuration belong in required reporting rather than in implementation footnotes.

Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization

Shiyun Wa, Yifei Wang, Anna G. Green, Simone Sciabola, Ye Wang Steering molecular generators toward desired properties usually relies on policy-gradient reinforcement learning, which needs a trajectory log-probability whose form depends on the specific architecture and generation procedure, so optimizers do not port across model families. EW-SFT (Elite-Weighted Supervised Fine-tuning) instead uses the reward only to select an elite set of high-scoring molecules, then updates the model with its own pretraining loss on that set; ablations indicate the reward signal flows mainly through elite selection rather than continuous weighting. Because the rule needs only scored molecules and the model's native loss, the same optimizer works across autoregressive, masked-diffusion, and discrete-flow generators and across de novo, motif-extension, and linker-design tasks, outperforming each model's native optimizer under a fixed budget of 3D shape alignment oracle calls on two kinase references.

Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models

Linhai Ma, Rita El Hachem, Mahatab El Hajj, Lilian Ghandour, Samah Fodeh Crisis helplines assess suicide risk through structured interviews that are slow and depend on operator training, and almost no prior work covers Arabic-language calls or works within real helpline privacy constraints. Using de-identified transcripts from Lebanon's National Lifeline — transcribed on site with a Levantine Arabic speech recognition model and scrubbed locally by an Arabic named-entity recognition model — the authors fine-tuned five instruction-tuned language models and six transformer encoder baselines on both the Arabic transcripts and machine-translated English versions, labeling calls with two binary outcomes derived from the Columbia Suicide Severity Rating Scale. Across 383 calls, the best English model reached a macro-F1 of 85.00 and ROC-AUC of 92.59 on high-risk classification, catching 88.9% of high-risk calls, with the best Arabic model close behind at 81.19 macro-F1; lower-severity ideation proved considerably harder in both languages.

CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships

Jacy Reese Anthis, Mark D\'iaz, Renee Shelby cross-listed Research on AI companionship is limited by scarce and unreliable human-chatbot interaction data. CompanionSim scales a small amount of real data into 2,240 simulated multi-turn conversations covering 16 chatbot behaviors across seven use cases, which human annotators then rated alongside real conversations in two studies (a U.S.-representative sample of 628, and 3,646 participants across the U.S., U.K., India, and Nigeria). Companionship behaviors such as validation reduced likability, humanlikeness, and trust rather than increasing them, with the effect strongest among women and older participants.

NSIDDx: A Design Framework for Neuro-Symbolic, Practitioner-First Differential Diagnosis in Low-Resource Settings

Aarav Singh Diagnostic pipelines built on a language model plus retrieval over rare-disease knowledge score well on benchmarks, yet evaluation on uncommon presentations across two cohorts showed they emit confident outputs that clinicians often cannot verify and that resist interrogation. NSIDDx responds with a design framework treating the clinician as an active reasoning agent rather than a recipient of answers, instantiated as a neuro-symbolic pipeline with ternary symptom encoding, contradiction detection, audit strings, and practitioner override that runs offline on consumer hardware. The authors distill five design principles for clinician-in-the-loop clinical language processing and call for prospective studies, positioning the work as a framework proposal rather than a validated system.

Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching

Kartik Ravisankar, Hojat Abdolanezhad, Daniel Capo, Sang Su Lee, Shishir Dash, Vijay Anand Raghavan As service marketplaces shift from fixed request forms to language-model matching that infers intent from free text, the provider-side attribute taxonomy underlying matching, search, and pricing has to be rebuilt in a form providers still understand. The described autoresearch loop generates that taxonomy one occupation at a time through iterative propose-evaluate-keep cycles, scoring each candidate tag set with a recalibrated six-rubric model-as-judge and applying weighted penalties from a seven-critic persona panel with no hard vetoes. A separate parity-mapping stage infers which provider attribute each legacy form question was meant to measure and maps it onto the generated tags, giving both a coverage signal and a human quality-assurance interface. The system has run in production at a major U.S. consumer services marketplace since April 2026 across 132 occupations.

Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries

Phuong Anh Nguyen, Jill Noorily, Matthew Flathers, Haruka Notsu, Laura Ospina-Pinillos, Tommy Nguyen et al. cross-listed As health information seeking moves from ranked link lists to conversational answers, the job of judging sources shifts from the user to the platform, yet little is known about what these systems actually cite. An audit of ChatGPT, Perplexity, and Google AI Overview on twenty English mental health questions under two prompt conditions, plus a three-question subset translated into six additional languages, recorded 15,942 citations across 1,140 responses and 1,713 unique domains, each classified by a validated nine-category typology. Citations were highly concentrated: the ten most-cited domains accounted for 43.6% of English citations, with government, commercial health, and academic sources each near 22%, and explicitly asking for sources changed the mix only modestly. Non-English queries returned fewer citations and were routed to language-appropriate resources significantly less often; the typology, classifier, and annotated corpus are released for reuse.

Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts

Saman Rahbar, Xiliang Zhu, Irvin Cardoza, David Rossouw Real-time agent-assist tools in contact centers must decide, for each of many predefined topics, whether a live customer utterance is relevant, working from ASR (Automatic Speech Recognition) transcripts of spontaneous phone calls that are unclear, repetitive, and largely unpunctuated. The authors curate a human-annotated dataset of topic-utterance judgments from real call-center transcripts and compare a regex baseline, zero-shot sentence embedding encoders, and Gemini-based matchers, crossed with two ways of describing a topic: keyphrases versus natural language descriptions. Lightweight LLM matchers paired with natural language topic descriptions outperform both embedding and regex approaches, indicating that how a topic is expressed matters alongside the matcher itself.

Hidden relationships in a document-derived property graph: top-k chunk embeddings and inverse-distance weighting over a dynamically evolving ontology

Bilge Kaan Karamete, Hunter Casten cross-listed Knowledge graphs that large language models extract from text capture only explicitly stated facts, leaving semantically related entities disconnected across documents. The proposed additive second pass leaves those facts untouched: each document is chunked and embedded once, top-k nearest-neighbour queries over existing chunks yield candidate node pairs via entity membership maps, and pairs are scored with Shepard inverse-distance weighting over a rescaled chord distance, avoiding the threshold collapse of affine cosine scoring behind a k-NN gate. Un-gated per-pair accumulators form a commutative monoid, making the pipeline strictly order-independent and incrementally scalable without recomputing earlier documents. Implemented across FalkorDB, Kinetica, ArangoDB, and Neo4j, 768- and 240-dimensional embeddings retain 92% and 72% edge fidelity against a 3072-dimensional baseline while the top-k formulation runs 25x faster.

Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations

Fanyou Wu, Suraj Maharjan, Ainur Yessenalina, Dennis Xu Chen, Rahul Srivastava, Srinivasan H. Sengamedu Managers heading into hard conversations with employees need to practice speaking aloud, which text chatbots cannot offer despite being the scalable option. Conversation Coach is a voice-first rehearsal system built around three requirements: low-latency interaction with strong language understanding, configurable bot personalities that simulate different employee types, and personalized feedback on content and policy compliance. Comparing an end-to-end speech-to-speech model against a cascade of automatic speech recognition, a large language model, and text-to-speech, the end-to-end route delivered 3x lower median (P50) latency with native barge-in at an estimated 8x lower cost, while the cascade reasoned well enough to matter for coaching quality and was the architecture actually deployed. In production it was used by more than 40,000 managers over six months, with adoption concentrated on genuinely difficult conversations.

EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models

Muran Yu, Jiechao Gao, Yuandong Pan, Barney H. Miao, Andrew C. Lesh, Kincho H. Law et al. Locally run small language models are appealing for scientific question answering because of privacy and deployment stability, but they face small literature collections, fragmented evidence, short context windows, and weaker reasoning. EGT-KG (Evidence-Grounded Typed Knowledge Graph) is a retrieval framework built to work under those constraints, compared against plain retrieval-augmented generation in two variants, one with an automatically generated relation schema and one with an expert-defined schema. Scored on a six-dimensional rubric covering soundness, correctness, completeness, conciseness, relevance, and fluency over a biopolymer-bound soil composite literature benchmark, EGT-KG beats vanilla RAG in most settings, with the largest gain on llama3:8b at a final score of 70.37, up 14.67%.

Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing

Alexandre Clin Deffarges, Nataliya Kosmyna, Pattie Maes A study of 50 participants learning nuclear safety protocols compared three AI tutoring designs: an unrestricted ChatGPT-style chatbot, a Socratic bot that gives hints but withholds answers, and a non-conversational tutor that adapts difficulty from Muse headband EEG signals measuring cognitive engagement. The unrestricted chatbot produced the highest immediate post-test learning gains (p < .03, d > 0.80), while the adaptive EEG-driven condition generated the highest measured brain engagement (p = .018). Clustering of interaction logs showed unrestricted-mode users mostly retrieved answers directly, whereas Socratic-mode users started reasoning through hints and then progressively disengaged. The authors argue the unrestricted chatbot's advantage reflects the immediate timing of the test rather than deeper learning.

CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN

Pranshav Gajjar, Vijay K Shah Telecom small language models (SLMs) embedded in AI-native 6G radio access networks typically justify decisions after the fact, and attempts to train pre-hoc reasoning traces with Group Relative Policy Optimization (GRPO) hit a cold-start barrier where models learn either the output format or the correct label but not both. CRAFT (Cold-start Reasoning Alignment via Fine-Tuning) sidesteps this by autonomously generating a verified dataset of input–trace–label triplets and fine-tuning with low-rank adaptation (LoRA). On the TRACTOR and IC xApp datasets it reaches up to 86.5% accuracy and 94.6% F1 with zero parse failures, while direct GRPO and supervised-fine-tuning-then-GRPO stay below 53.5% F1, and it uses 59% less energy; CRAFT-initialized policies also remain stable under subsequent GRPO training with varied reward functions.

TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data

Jianhuai Hu, Yashan Wang, Shangda Wu, Zhancheng Guo, Shijie Liang, Wuna Meng et al. cross-listed Audio-to-score transcription models are limited by the scarcity of paired audio and notation data, which confines most systems to a single instrumentation. TUTTI drops real scores entirely for pre-training: a symbolic music generation model produces a large multi-instrumentation corpus, which is rendered into audio-score pairs with expressive acoustic variation and used to pre-train a standard Transformer encoder-decoder. Pre-training on synthetic multi-instrumentation data yields stronger representations than single-instrumentation training, and after fine-tuning on real datasets the model sets new state-of-the-art results across audio-to-score baselines and transfers competitively to instruments never seen in training; code and the TuttiCorpus dataset are to be released.

SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification

Swapnil Bhattacharyya, Mayank Baranwal When a language model rewrites a mathematical optimization problem into another modeling language, checking the result by running a solver is unreliable, since local minima, timeouts, and numerical artifacts can mask genuine semantic divergence. SOVER splits the job in two: the model extracts a variable mapping between formulations, then Z3 formally certifies domain cross-feasibility and objective-order preservation for mixed-integer linear problems while dReal supplies tolerance-aware feasibility, range, and approximate-argmin checks for continuous nonlinear ones. On NLEquiv-150, a new public benchmark of 100 equivalent and 50 deliberately hard non-equivalent nonlinear reformulation pairs, the system classifies 149 of 150 pairs correctly including all 50 hard negatives, with the single failure traced to an incomplete mapping extraction.

AnalysisBank: An Expert Analysis Pattern Library for Financial Report Generation

Yajing Yang, Yunshan Ma, Kelvin J. L. Koa, Min-Yen Kan Automated financial report generation typically plans at the structural level, deciding which topics or sections to include, which tends to produce content that restates rather than analyses. AnalysisBank distils expert reports into a library of Analyses, each binding a data signal to an analytical move and the expert text span it came from, then at inference matches the input's signals against the library and applies the retrieved moves. Distilling 550 expert reports yields a heavy-tailed distribution of 47 to 52 signal types across 13 move types, and on two financial benchmarks with four LLM backbones the approach raises the share of novel, data-grounded insights by 1.7 to 3.7 times over structural-level baselines, with transfer experiments on scientific writing suggesting the analytical-versus-structural distinction is not finance-specific.

Staged Linguistic Seeding: Grounded Query Expansion for Verified-Unit QA in AI Contact Centers

Hyeonseop Yoon, Jeong-Eun Park Customer-service question answering in a voice contact center faces latency limits and a high cost for wrong or unsupported answers, so the deployed system answers only from a closed set of human-verified units, returning one verbatim or routing to clarification, abstention, or a human handoff. Coverage comes from staged linguistic seeding, an offline index-enrichment step where a human writes a per-unit grounded slot recipe, gpt-4.1-mini renders it into query variants, and a light human gate filters them, leaving inference as a single retrieval pass with no query-time generation. On held-out variants from two industrial domains, recall@1 rises to 0.881 and 0.930 (gains of 0.27 and 0.34) with improvements across all five retrievers tested, beating doc2query at the same generation budget, and the verified-unit design cuts unsupported content from 7-13% to roughly zero.

Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair

Anik Jha cross-listed When a code model's first attempt fails a test, fault localization is supposed to help by pointing the repair at the implicated statements, but a targeted edit may simply benefit from being small, and a second sampling call may fix things without using the failure at all. Three arms applied to the same failed candidates — blind whole-solution resampling, spectrum-based localization with suspect-span infilling, and same-length infilling at a random disjoint span — were run across three frozen 26-32B models, three benchmarks, 488 failing candidates, and a separately declared 24B fourth model. Localization was usable on only 9.0% of failures, and among the 177 localizable candidates localized infilling lost decisively to blind resampling at matched attempt counts (3:40, p = 3.0e-9), a result that replicated in a third model family; infilling reproduced the removed span verbatim 48.9% of the time, explaining why extra budget does not help, and the advantage over the random-span placebo held only pooled, not per model.

Conditional Flow Matching for ML-Based Inverse Design Problems

Juliana Felder, Milad Habibi, Soheyl Massoudi, Mark Fuge Engineering inverse design is slowed by iterative solvers for problems constrained by partial differential equations and by their sensitivity to where optimization starts, which motivates generative models that propose candidate designs without rerunning the simulator. Conditional flow matching is added to EngiOpt and compared against a conditional diffusion model and a conditional generative adversarial network on the beams2d structural and heatconduction2d thermal tasks from EngiBench, scoring generated designs as warm starts for gradient refinement via cumulative and final optimality gap. Flow matching achieves the lowest measured optimality gaps, maximum mean discrepancy, and volume-fraction deviation on both tasks, and at 16 Euler steps reaches 53.2 samples per second on beams2d, roughly 66 times the throughput of the diffusion baseline at 1000 network evaluations.

A Dataset for Modeling Iterative Problem-Solving

Fagun Patel, Sang T. Truong, Duc Q. Nguyen, Kazunori Fukuhara, Benjamin W. Domingue, Sanmi Koyejo et al. Solving a problem through repeated attempts is a sequential process in which a solver receives feedback and revises, and predicting whether performance improves, plateaus, or regresses matters for understanding both human learners and autonomous agents. CodeInsight captures this at scale with over 3 million submissions from 3,286 undergraduates in two introductory C++ courses over two academic years, including test-case-level outcomes, timestamps, and source code. A benchmark on the dataset compares parametric, sequential, and generative predictors under a shared calibration-and-scoring protocol, and a Recurrent State Space Model (RSSM) adapted to track solver traits through discrete latent variables is the most accurate on three of four courses. An LLM-based predictor that writes full submissions is less accurate but exposes failure modes, and its coding proficiency turns out to be inversely related to predictive performance, making it better understood as a generative solver than a faithful model of student behavior.

ClinTraceBench: Source-Verifiable Longitudinal Clinical Reasoning over EHR-Derived Dialogues

Huimin Wang, Zhengyi Zhao, Yutian Zhao Clinical assistants built on large language models (LLMs) must reason over multi-visit patient histories, but whether compact history representations such as retrieval, summaries, or agentic memory preserve the needed longitudinal signal has not been measured. ClinTraceBench provides 385 verified dialogues derived from MIMIC-IV with event-level provenance and a nine-task taxonomy, used to evaluate eight history strategies, including full context, BGE-M3 dense retrieval, LLM summaries, and the agentic memory systems Mem0 and A-Mem, across DeepSeek-V3, GPT-4o-mini, Haiku 4.5, and Sonnet 4.6 on 6,271 questions. In a controlled injection probe, Mem0, A-Mem, and LLM summarization recover only 0 to 5.3% of injected attribution facts even when the sentence is present before memory construction, and compressed strategies pay an aggregation tax on multi-visit trends and cross-patient comparisons. The gap between no context and full context ranges from about 30 to 63 percentage points depending on backbone, and on the cost-accuracy frontier Haiku 4.5 with full context dominates Sonnet 4.6 at roughly a quarter of the cost.

When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting

Takumi Fujimoto, Hiroaki Nishi Online adaptation can help time-series forecasting on edge devices under distribution drift, but the measured benefit depends heavily on evaluation choices. Using six public multivariate streams including building-sensor and smart-meter data under a leakage-free streaming protocol, the study identifies two sources of comparison bias: the static baseline's warmup budget, which shifts the estimated adaptation benefit by 3.0 to 18.8 percentage points across a 1,000 to 20,000 step range, and comparing SGD with momentum against Adam at a shared default learning rate, which conflates optimizer quality with rate sensitivity. When warmup and per-optimizer learning rates are selected on a held-out pre-drift validation slice, Adam outperforms SGD with momentum in 310 of 360 evaluated cells, though four Adam cells remain below the static baseline. Measurements of adaptation-state memory and A100 per-update latency on PatchTST show several parameter-efficient variants are nondominated on the memory axis, and reported smart-meter gains depend on meter-selection rules.

Athena: Vulnerability-Affected Library Identification via Knowledge Graph Completion

Phong Trinh Duy, Trang Dang Yen, Hung Nguyen-Huu, Bach Le, Quyet-Thang Huynh, Dieu Hoang Vu et al. cross-listed More than half of vulnerability database entries have missing or incorrect affected-library information, and existing automated approaches treat identification as isolated text retrieval, ignoring the relational structure of the databases. Athena models vulnerability databases as a knowledge graph integrating CVEs, libraries, CWE weakness types, CPE products, and software ecosystems, reformulates the problem as knowledge graph completion (KGC) via link prediction, and re-ranks KGC candidates with a fine-tuned LLM augmented with knowledge graph embeddings. On VulLib, Athena improves average F1 by 32% over the best of four baselines, VulLibGen, and its 110M-parameter KGC backbone alone already surpasses VulLibGen's best 7B-parameter configuration, with re-ranking adding consistent further gains across all evaluated LLM backbones.

Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment

Yin Fang, Qiao Jin, Shubo Tian, Lauren He, Maya Geer, Noor Naffakh et al. Many cancer clinical trials fail from insufficient enrollment, and existing AI recruitment tools mostly assess eligibility in isolation without evaluation in real oncology workflows. TrialGPT 2.0 is a trial recommendation system that also judges which trials warrant consideration given a patient's current clinical needs and local workflow priorities, and it provides structured, inspectable explanations for expert review. In retrospective multicenter cohorts of 288 cases it surfaced at least one clinician-recommended trial in its top 10 for about 91% of cases while cutting clinician screening time by 55%, and a six-month prospective deployment in an active precision oncology tumor board expanded patient access to trial participation by 90.9% by catching opportunities the routine workflow missed. The authors also release NIH-TrialBench, 126 clinician-authored synthetic patient vignettes and matching scenarios from 11 NIH Institutes and Centers.

Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening

Zhilong Song, Lixue Cheng cross-listed Crystal generators and tool-using agents now propose candidate structures faster than density functional theory (DFT) energy and phonon calculations can vet them, yet most screens check little beyond atomic overlap and give no chemical reason for rejection. Here agents generate, test, and actively refute two million candidate laws, distilling eight Plausibility Rules for Inorganic Structures (PRIS) that encode short-range repulsion, ionic contact and packing, electrostatic balance, bond-valence conservation, and crystallographic site complexity. Experimental structures satisfy the rule sets at 82 to 99% versus 6.5% for Pauling's rules 2 through 5 combined, and the strictest set detects 87.9% of damaged crystal structures where distance cutoffs catch only 1.6 to 3.2%. A derived synthesis score screens 83.7% of hard-to-synthesize structures while retaining 80.7% of experimental ones, cuts a DFT validation queue by up to 67.3% in an inverse-design run, and explains why GNoME is enriched in rare low-symmetry structures.

From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs

Jie Chen, Xiangqian Yu, Yanchao Lian, Tan Lu, Run Yang, Zhengchun Shang et al. cross-listed Scaling Transformers over user behavior sequences for production recommendation ranking faces two obstacles absent from language modeling: behavior signals are noisy, temporally irregular, and sparsely supervised, and each request must score many candidates against one shared user history under tight latency. ReST addresses signal quality with a sequence encoder using dual-gated attention, rotary positional and temporal embeddings, stabilized residual normalization, and training-only auxiliary objectives, and addresses compute asymmetry by splitting ranking into a heavy reusable encoder and a lightweight cross decoder with projection-free key-value attention, so a user prefix is encoded once and decoded many times. Across industrial and public benchmarks it scales more consistently along sequence length, depth, and width where LLM-style blocks saturate, and a one-week online A/B test on a production advertising platform raised online AUC by 1.31% and a core revenue metric by 11.93% within a 50 ms P99 budget, after which it was fully deployed.

Relational-Core Graph Analytics Querying graphs at SQL scale, and why the node/edge model is a performance tax, not a truer picture of connected data

Gene Zhang cross-listed A durable assumption holds that graph analytics needs a purpose-built graph engine and that relational systems handle connected data badly; the argument here is the reverse for the workloads enterprises actually run. ClickGraph and its Databricks-dialect sibling DeltaGraph translate Cypher directly onto existing relational schemas — the tables, columns, and foreign keys as they already stand — executing in place on ClickHouse, Databricks, or lakehouse files with no import and no separate cluster, which leaves ordinary SQL as an open optimization surface. The authors contend the node/edge property graph is a re-encoding of relationships relational tables already store explicitly, making query-time reconstruction pure overhead, and support this with a peer system's own published benchmark in which a columnar engine outruns Neo4j by two to four orders of magnitude plus reproducible runs across the LDBC Social Network Benchmark suite.

Can LLMs Design Video Coding Tools? A Case Study on Planar Mode

Yingwen Zhang, Meng Wang, Liqiang He, Shiqi Wang cross-listed Designing video coding tools resists automation because any tool change couples tightly to the rest of the codec; the case study asks whether an LLM can redesign Planar mode, a long-standing intra prediction tool in video coding standards. The setup is a generation-and-evaluation loop in which the model proposes Planar predictors, encoder trials measure coding performance, and the model revises from that feedback. Replacing the default Planar mode in the Fraunhofer Versatile Video Encoder (VVenC) under its faster preset, the generated version achieves 0.18% bitrate savings for 0.4% complexity overhead on the standard benchmark. Extending to the Enhanced Compression Model (ECM), both replacing the newer directional Planar modes and adding the generated predictor as an extra mode with its own syntax elements produced gains in a constrained low-resolution setting.

StudentSim: Training LLM-based Student Simulators

Ke Yang, Chenglong Wang, Michel Galley, Chandan Singh, Jeevana Priya Inala, ChengXiang Zhai et al. AI tutors work best when they adapt to individual students, but evidence about which guidance suits which learner is slow and costly to gather, and existing student simulators either track state without processing explanations or role-play fluently without matching the target student's competence. StudentSim turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization, so each simulator both mirrors a student's own responses and updates them under tutor guidance. The accompanying StudentSimEval protocol covers 60 students across chess, second-language English writing, and mathematics and measures behavioral fidelity and guidance responsiveness; StudentSim outperforms GPT-5.4 on both metrics in all three domains, reaching fidelity 0.51 and responsiveness 0.91 in chess versus 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. Used as a reward model for tutor reinforcement learning, it produced a chess tutor that expert humans rated more accurate, better-guided, and more personalized than both a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward.
68 more specialized papers

Agents 82

HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models

Yun-Jian Zhang, Chen-Wei Liang, Tian-Yi Zhang, Jian Ding, Yi-Lun Wu, Ao-Bo Li et al. Text-based world models must learn symbolic action effects from serialized state descriptions, but how that state is formatted has been largely unexamined. HyperWorld compares raw observations against three symbolic serializations of identical ground-truth state — independent sentences, pairwise triples, and entity-centered hyperedge units that bundle several related facts around an entity or relation — under one training objective that predicts effects or flags an action infeasible. Hyperedge grouping helps most at 0.5B–1.5B parameters and under distribution shift, taking the best out-of-distribution fact F1 and the highest success rate in downstream greedy planning, while larger models narrow the gap and pairwise triples can edge ahead on in-distribution exact match.

Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls

Dheeraj Mohandas Pai, Lu Xian Long-horizon agent benchmarks report end-to-end success but mix state-tracking difficulty with instruction ambiguity and can be gamed by a hallucinated final answer, leaving it unclear whether a model can carry exact intermediate state at all. The test here is computing an MD5 hash step by step: 196 dependent tool calls over 64 rounds, with the model holding four 32-bit words in its own context between calls, checked against a from-scratch RFC 1321 reference trace so any error is pure bookkeeping. gpt-oss-120b, a mixture-of-experts model with only about 5.5B active parameters per token, carries state across all 196 calls and returns the correct digest on a majority of completed runs, including a variant where every arithmetic primitive is replaced by a second LLM worker; success hinges on keeping the model's own reasoning in context each turn and voting over a thinking-enabled worker to cancel modular-arithmetic slips.

OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets

Dongsheng Chen, Xiangyu Zhao, Xin Yao, Xuetao Wei As deployments grow into fleets where several agents, planners, and execution backends touch the same environment, safety becomes a question of whether a concrete action should be allowed to commit rather than whether a prompt looked risky. OpenAgentFlow splits the system into a control plane and an action plane, normalizing pending GUI actions, API calls, tool calls, and LLM-generated invocations into one AgentEvent stream that passes through a shared pre-execution policy enforcement point, with provenance, session state, audit records, and hot-updatable policies held centrally so new rules apply without touching agents, prompts, or models. An Android instantiation reaches 94.0% accuracy and a 95.3% attack block rate on a 300-case action-event benchmark, matches expected behavior in 27 of 30 dynamic-policy cases after new rules are installed, and scores a 92.9% trace-adjusted pass rate on emulator traces spanning GUI, API, and planned actions.

UI-Venus-2 Technical Report

Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao et al. Graphical user interface agents tend to be tuned for benchmarks rather than deployment, held back by narrow environment coverage, brittle task construction, and reward signals that cannot be trusted. UI-Venus-2 scales three axes together for a single closed-loop reasoning-and-action foundation agent spanning mobile, web, and desktop: coverage of more than 170 multilingual mobile apps plus native desktop operating systems, a deep-research pipeline that generates function-grounded instructions, and trace-level plus sample-level verifiers using visual keypoints and multi-model voting to produce reliable reinforcement learning rewards. Safety-aware mechanisms gate consequential actions, and the model is released open source.

EULER: Exploring Underused Links with Evidence-Checked Return for Multi-Agent Mathematical Discovery

Ren Zhenzhuo Moving a mathematical problem between communities that use different objects, invariants, and tools is costly, so such transfers are usually skipped. EULER makes that transfer — a bridge — the unit of search in a multi-agent system: direct, adjacent-domain, and distant-domain routes compete for budget around a fixed conjecture, and a bridge keeps its budget only if it supplies an operation the source representation cannot execute and its target-side evidence returns to the original statement through a checked implication, with six ordered stress tests filtering invalid bridges before expensive search. On 120 recent combinatorics conjectures frozen and screened for contamination, the system produced 10 proofs, 3 refutations, and 45 scoped partial results; ablations show the stress tests cut incorrect conclusions from 9 to 3, and success tracked executable operation gain and valid return rather than domain distance.

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

Hadi Mohammadi Production evaluation of LLM agents usually shows a judge only the request and the final reply, which is structurally incapable of noticing an agent that got the right answer the wrong way. The setup measures that blind spot with ground truth by construction — a deterministic tool-using support-desk environment, a scripted oracle policy, and a fault injector that breaks exactly one thing at a known step, with faults stratified by whether the customer-visible outcome survived. Across 400 trajectories and five judges, the outcome-only judge catches 84% of outcome-breaking faults but only 45% of silent ones while falsely flagging 33% of correct trajectories, whereas a step-rubric judge reaches 77% silent recall with zero false alarms at three times the cost; no judge reads the final reply, so an invented promise appended to a perfect trajectory slips past the rules entirely and the step judge 82% of the time, and self-consistency triples cost without improving anything.

GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments

Lin Fu, Zheyuan Yang, Tianhui Zhang, Jinbiao Wei, Guo Gan, Boxu Liu et al. Models that generate the next graphical user interface screen are usually scored one step at a time, even though their purpose is to serve as multi-step environments where generated states get reused as the basis for further interaction. GUI-CC tests that reuse directly with two tracks: an offline track rolling models along 500 real mobile trajectories drawn from GUIOdyssey, and an online track where fixed probing agents interact with model-generated interfaces on 200 emulator-verified tasks across 30 apps, scoring transition fidelity, transition plausibility, contextual consistency, and task progress. Plausible single-step generation turns out not to imply reliable simulation — current models render usable-looking screens while losing task-relevant context and failing to support executable multi-step rollouts.

Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness

Sagar Srinivas Sakhinana, Venkataramana Runkana cross-listed Autonomous agents that modify infrastructure, deploy services, and verify their own results need explicit machinery for progressing through long tasks, containing what they can execute, and recovering from failure. The proposed framework separates three concerns: graph engineering encodes workflow progression with verification-gated transitions, loop engineering bounds diagnosis, repair or re-planning, retry, and re-verification, and an agent harness enforces zero-trust execution through identity, authorization, policy-scoped capabilities, isolation, and runtime safeguards. Instantiated on Google Cloud, every run terminates either in a verified operational deployment or an auditable terminal failure within its recovery bounds, with progression requiring machine-checkable repository, deployment, and runtime evidence.

AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes

Xun Wang, Bihe Zhao, Michael Backes, Franziska Boenisch, Adam Dziedzic cross-listed Commercial large language model (LLM) APIs may silently substitute, quantize, or wrap the backbone they advertise, and every existing audit infers identity from the text channel — which agentic serving stacks discard once the model emits a tool call, and which provider-injected system prompts can distort enough to falsely accuse honest providers. AgentProv instead fingerprints a deployed model by its categorical tool-call distribution and decides identity with a maximum mean discrepancy (MMD) permutation test, on the observation that agentic post-training writes tool-use behaviour into the weights in a way that survives deployment context. It caught every substituted model across 630 evaluated checkpoint pairs, a 100% detection rate, while holding the false-positive rate under system-prompt injection to 7% versus 67% for MET and 53% for RUT; on third-party endpoints its disagreements with MET line up with a token-count side channel that reveals injected system prompts.

CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language

Qi Fan, An Zou, Yehan Ma Writing fast CUDA kernels requires algorithm design, correctness validation, and hardware-aware tuning, and prior LLM work mostly transpiles from PyTorch to CUDA rather than generating kernels directly from natural language (Text2CUDA), where the model must bridge high-level intent and low-level implementation. CUDA-Harness introduces Intermediate-Structured Generation to connect semantic understanding with kernel code, Synthesis-Based Verification that supplies isolated test data and progressive validation to blunt the reward hacking that comes from relying on predefined test inputs, and Feedback-Adaptive Evolution that prioritizes correctness before optimizing speed. Experiments report gains over prior approaches along with generalization across different LLMs, hardware platforms, and C-to-CUDA transpilation.

A Formal Analysis of Agent Payment Protocols

Ke Jiang, Mohan Yu, Yuan Chang, Mohit Kumar Jangid, Jianyu Niu, Cong Wang et al. cross-listed Agent payment protocols let AI agents buy goods and settle payments for users, spreading intent, delegated authority, credential use, settlement, and fulfilment across actors and stages in ways that no single message can enforce, yet their security guarantees remain implicit across specs and reference implementations. Four representative protocols — x402, MPP, ACP, and AP2 — are modelled in the Tamarin prover under a shared abstraction of the payment lifecycle, using source-backed verification questions and counterexample traces rather than a presupposed property taxonomy, yielding 18 shared security principles. Across 86 verification cases the analysis reproduces 46 known or calibration cases and surfaces 40 previously undocumented formal-consistency findings, each traced to a missing protocol relation and reverified against a minimally strengthened model, with ten validated through proof-of-concept exploits, schema-level witnesses, and executable traces.

Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents

Timothy Kassis, Vinayak Agarwal, Yuhuan He, Darshil Patel, Aubrey M. Brueckner An agent asked to analyse an experiment will usually produce running code, but whether the analysis is defensible depends on procedural choices — which statistical test the field accepts, which identifier namespace is authoritative, which caveats must accompany a result. Scientific Agent Skills is an openly licensed library of 163 such procedures across 16 areas of practice, spanning genomics, cheminformatics, medical imaging, study design, and scientific communication. Each skill is a directory around a versioned human-readable instruction file that the agent loads only when a task calls for it, often alongside reference material and runnable scripts; the authors report no task-level evaluation and no measurement of how often hosts select a skill.

AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis

Yuetong Wu, Maojun Sun cross-listed Powder X-ray diffraction (XRD) analysis resists automation because an agent must read diffraction evidence, drive refinement software, order coupled parameters defensibly, and tell a numerical improvement from a physically valid one. AutoXRD structures the task as stepwise refinement grounded in observed evidence with deterministic crystallographic and physical checks gating acceptance, and XRDBench evaluates it on 100 bounded reasoning tasks plus 34 executable end-to-end workflows requiring file inspection, software execution, iterative refinement, and reporting. Across 1,340 model-task runs, ten recent LLMs average only 57.8 out of 100, dropping from 61.9 on the reasoning track to 53.7 end-to-end, with GPT-5.6 Sol highest overall at 81.1; execution traces expose recurring failures in coupled-parameter control, quantitative reasoning, evidence preservation, and knowing when to stop.

Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops

Bowei He, Weixu Zhang, Yili Jin, Xue Liu Code-level autonomous research loops, where a large language model (LLM) agent proposes edits to a training pipeline, runs it, and keeps changes that improve a measurable in-loop metric, are examined here for whether those gains reflect genuine progress. Across several experimental settings the authors identify a failure they call algorithmic mode collapse: edits keep touching different lines of code while the underlying kinds of algorithmic change repeat, and in-loop gains diverge increasingly from held-out evaluations. Their mitigation, DAPS (Diversity-Aware Proposal Sampling), combines category-coverage reweighting, a persistent edit memory, and a validation gate, and under a three-tier protocol separating in-loop, audit, and blind metrics it cuts semantic-cluster decay of edits by 69.1% and improves relative faithfulness by 83.7% on the blind metric while preserving in-loop optimization speed.

Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations

Ao Qu, Panagiotis Michelakis, Linyuan Han, Yiannis Hadjiyianni, Kun Ouyang, Konstantinos Siskos et al. FAIRY is a full-stack agent system deployed on an operating soybean research farm at Harbin Institute of Technology, covering a whole season from ridge preparation and planting through irrigation, fertilization, pest and disease treatment, harvest, grain handling, drying, and storage. It integrates production machinery APIs, fixed soil and canopy sensors, multispectral and thermal drones, satellite vegetation products, a weather station, calibrated crop-process models, and multi-season yield histories under an "everything is an event" execution paradigm, on top of which sit a library of atomic agronomic skills, multi-agent controllers, frontier and edge model execution, and full-path trace logging. The authors evaluate nine state-of-the-art agent controllers across one hundred full-season scenarios on a 64-ridge field, scoring agentic success, spatiotemporal correctness of the entire action path, token cost, and edge-device runtime.

ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation

Muzhao Tian, Zezi Zeng, Yifan Yang, Xin Gao, Yan Li, Zisu Huang et al. Slide generation from documents demands both faithful content selection and precise spatial layout, but current slide agents rewrite a whole slide or deck, render it, and only then critique — delayed feedback that makes local failures like overflow, overlap, clipping, and off-canvas placement hard to attribute or repair. ReDeck breaks revision into atomic edit actions and returns renderer-derived observations after each one, turning the loop into "one edit, one observation," layered with turn-level adaptive critique for semantic and design guidance and a submission-level gate enforcing hard layout validation. Paired with DeckQuiz, a benchmark separating content fidelity, spatial correctness, and design quality, ReDeck outperforms existing slide agents across GPT-5.4, Claude-4.6, and Gemini-3.1, with ablations showing feedback timing and granularity both matter.

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

Haechan Kim, Yoonho Lee, Gisang Lee, Chelsea Finn, Kangwook Lee An agent's performance depends on both its model weights and the executable harness that manages context and control flow, and tuning either alone leaves the system bottlenecked by the frozen half — weight updates change which harness works, and harness changes expose different model capabilities. WHALE (Weight-Harness Alternating LEarning) alternates two phases, updating weights under the current harness via online rejection-sampling fine-tuning, then searching for a better harness under the updated model via Meta-Harness, switching on either fixed phase durations or an adaptive patience rule. With Qwen3.5-2B/4B agents on search question answering, mathematical reasoning, and chess puzzles, it beats weight-only, harness-only, and Fast-Slow Training by 4.15 to 24.38 percentage points in best mean@8 accuracy, and the bottleneck genuinely varies: harness search matches peak weight-only accuracy with far fewer rollouts on SearchQA, while math improves only after a weight update.

Don't Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field Changes

Pruthvi Davineni cross-listed When a language model agent proposes a fix in a GitOps workflow, applying it means editing a version-controlled config file, and the natural implementation — having the model write the diff or the new file — turns out to be unsafe for unattended automation on real Kubernetes manifests. Under strict patching almost no unified diffs apply, while a tolerant tool like GNU patch applies 96% but silently misapplies roughly 1 in 7 (14-20%) with no error signal; full-file rewrite is capability-dependent, with a small model corrupting files and a frontier model usually correct but nondeterministic and costing O(file size) per edit. The proposed alternative has the agent emit only a structured field-change intent — which resource, field, and value — while a deterministic pipeline indexes manifests by kind and name, locates the target scalar's exact character span through the YAML parser's node position marks, and replaces just that span in the raw text, preserving comments and formatting at O(1) generation cost. It ships as KubeAstra under Apache-2.0 with the benchmark released, scoped to faithfully applying a known change rather than judging whether the change is correct.

Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems

Rakibul Hasan Rajib, Mengxing Zheng, Qian Lou Orchestrating multi-agent language model systems requires deciding what each agent does next, and both obvious options fail: routing on the query alone cannot react to intermediate progress or errors, while routing on the full execution history forces every later decision to reprocess redundant steps, inflating cost. Gated-Memory Routing conditions each decision on the query plus a learned execution memory, where a Memory Write Gate commits only non-redundant reasoning steps, a Retrieval Gate hands each agent a compact relevant subset, and an Adaptive Halting Controller stops once the memory holds enough evidence to answer. Across five reasoning and code-generation benchmarks it achieves the best average accuracy, beating the strongest baseline by 2.44 points while cutting HumanEval inference cost by 31.9%.

Invalidation Contracts for Cross-Episode Agent Memory

Michael Wu, Arquimedes Canedo Agents that cache recovery advice from API errors save tokens across episodes, but server-side data changes silently turn those cached fixes wrong, and re-deriving every time erases the savings. Invalidation contracts attach version stamps and cacheability hints to each recovery suggestion so a client can evict exactly the stale entries, and they split realized savings into validity (fraction of cached fixes still correct after drift, a property of the protocol alone) and compliance (fraction the planner actually applies first try, a property of the model). Across seven models, three serving paths, and roughly 9,400 episodes, row-level invalidation raises compliance by up to 66.7 percentage points and recovers 29-33% of baseline token cost on four models, while table-level invalidation drops post-drift first-try rates to zero on five of seven. Identical wire bytes produced 100% first-try compliance on Claude Haiku 4.5 but 11% or below on Claude Sonnet 5, which refused fixes adding fields absent from the original request.

Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems

Panduranga Sai Varma Dantuluri, Jyotirmoy Sundi cross-listed When agents hold credentials, call services, and spawn sub-agents, the distributed-systems question of who may act on whose authority becomes acute because the deciding component is a language model an adversary can hijack. The authors argue for evaluating agent security under an untrusted-model assumption, where a fully prompt-injected agent still cannot exceed its explicitly delegated authority, and derive eight requirements from four adversaries: confused deputy, token theft and replay, prompt-injection privilege escalation, and compromised sub-agents. A default runtime using broad bearer credentials with authorization decided inside the model fails all four, and among LangGraph, CrewAI, AutoGen, and the Model Context Protocol authorization model, three offer no built-in confinement and one only partial. Their authorization broker blocks all four threats, accepted 0 of 200,000 forged tokens, confined a compromised sub-agent to a mean of 1.5 reachable actions versus all 8,100 under bearer delegation, and costs about 2.6 microseconds per decision.

The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems

Bardia Mohammadi, Laurent Bindschaedler Agent fleets now take actions that cannot be undone — moving money, deploying code, deleting data, disclosing information — and because current controls authorize one effect at a time, a set of individually correct decisions can still exhaust a principal's overall risk tolerance under a shared trigger. The proposed irreversibility budget makes residual value-at-risk a first-class resource that a trusted runtime accounts for per principal across agents, workflows, and tenants, charging each effect its residual loss and denying the marginal action once the aggregate would overdraw. In a controlled study, per-effect gates permitted fleet-level overdraws of up to 48 times the tenant's risk limit while the budget kept every correctly priced run inside it. The authors name conservative, dependency-aware pricing of heterogeneous, adversarially declared, correlated effects as the open problem blocking deployment.

Toward Workflow-Aware Benchmarking for Healthcare NLP Agents

Junyi Yao, Baichuan Li, Zihao Zheng, Jiayu Long Healthcare agent evaluations mostly use static medical question answering or one-shot generation, which omit longitudinal state, interruptions, and handoffs to humans. The proposed episode-level protocol separates evidence attributable to the model, the agent, and the simulated workflow; specifies a five-field episode schema; and defines annotation and scoring for state continuity, evidence traceability, and escalation decisions, with cost-sensitive treatment of missed versus unnecessary escalation. It is instantiated as four task templates — documentation update, evidence retrieval, patient messaging, and triage handoff — and is explicitly framed as a reproducible intermediate layer between static benchmarks and prospective workflow studies, not a measure of clinical outcomes.

Dr. Claw: An AI Scientist Workspace for Vibe Research

Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan et al. Command-line coding agents can already read and write files over long sessions, yet research work still scatters across chat tools, IDEs, terminals, and writing environments, and the decisions that would make it auditable go unrecorded. Dr. Claw is an open-source workspace that wraps existing coding-agent executors in a human-in-the-loop workflow rather than adding another autonomous agent, using persistent state objects, a reusable skill library, and multi-executor coordination to tie planning, execution, and writing into one traceable and recoverable loop. Evaluated against a bare command-line agent sharing the same backend executor, so the comparison isolates the orchestration layer, the wrapped system scores higher on research completeness while leaving an auditable, recoverable process trail; the paper also walks through an interactive three-view scenario and a failure-recovery case.

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

Maya Moriya, Sigal Raab, Yael Vinker, Tali Dekel cross-listed Origami knowledge circulates mostly as unstructured demonstration videos, while computational tools need structured representations such as crease patterns or executable parametric plans. FoldingAgent bridges the two with a vision-language model agent that infers explicit parametric folding programs from video, using specialized tools to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions over a parametric space of paper geometry plus folding actions. Because the agent acts sequentially and can re-plan, it mitigates the error compounding that plagues multi-step folding. On PurelandFold, a newly curated benchmark of Pureland origami videos with ground-truth geometry and action labels, the combination of model reasoning, tools, and physical simulation converts unstructured demonstrations into executable, physically plausible folding procedures.

RestoreBench: Can AI Agents Restore Power Flow Convergence?

Riccardo Mansutti, Andrea Pomarico, Robert Jakob, Qian Zhang, Alberto Berizzi, Kevin O'Sullivan Working out why a power flow case fails to converge and fixing it takes engineering judgment, experimentation, and iterative decisions inside a constrained action space — a plausible but largely untested target for tool-using LLM agents. The benchmark specifies the simulation environment, observation and action spaces, and evaluation metrics over two power grids with 46 non-convergent cases each, every case requiring one or more corrective actions to restore convergence. Several large language models are compared across three architectures — plain chatbot, single agent, and multi-agent — giving a reproducible starting point for agentic systems in power system planning and operation, with code released publicly.

SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation

Songwei Dong, Bingyan Lu, Makayla Kienlen, J. Nicholas Laneman, Cong Shen Spectrum management decisions increasingly require pulling together policy proceedings, legal regulations, and license databases that are disaggregated, mix text with tables, and are formatted for human readers rather than machines. SpecMind is a multi-agent retrieval-augmented generation (RAG) system where coordinating agents dispatch specialized sub-agents to retrieve and synthesize across these heterogeneous sources, evaluated on SpecBench, a new question-and-answer dataset built from real license records and policy proceedings to fill a gap in domain evaluation resources. The system reports over 80% win rate against strong general-purpose RAG baselines across spectrum tasks, which the authors attribute to more accurate retrieval and better contextual reasoning from the agent-based decomposition.

SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents

Rayan Khoury, Shih-Yao Lin, Pratyush Mishra Judging a task-oriented dialogue turn means checking whether it actually advanced the underlying workflow state, something holistic LLM judges can miss because they weigh the whole context at once and need at least one full model call per turn. SAGE compiles a workflow specification and per-turn state diff into atomic, schema-grounded criteria, then routes each through a cascade of symbolic rules and on-device encoder/natural-language-inference verifiers that abstain rather than guess, aggregating criterion verdicts into a turn-level decision with an evidence trace. Its recommended SAGE-Core operating point settles 81-91% of criteria at zero paid model cost, and across four slices of MultiWOZ, Schema-Guided Dialogue, and ABCD no LLM-as-a-judge baseline significantly beats it — including a state-aware GPT-4.1 judge costing $4.7-8.0 per 1,000 turns. A two-annotator audit (n=200, kappa 0.94) supports label fidelity on transcript-visible failure classes, while the authors scope construct-validity limits from injected failures and partial symbolic circularity.

mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers

Timothy Kassis Handing an agent a dossier about a named expert is often claimed to do three separate things at once: supply hard-to-find material, produce a recognizable persona, and improve the agent's judgment. mimeo is an open-source tool that gathers a person's public work, verifies each extracted quotation against cached source text (rejecting 13.2% of them), and emits a loadable agent file, tested here across four expert profiles in one coding-agent harness. Knowledge access was the clearest win, with mimeo answering all 20 obscure quotation-heavy questions while no closed-book condition exceeded 10, though BM25 keyword search over the same pages answered 15-17; grounding also prevented the factual misstatements that memory-written personas made on 1-4 of 20 answers. Judgment transfer stayed unresolved because every condition hit the ceiling on the engineering and application tasks, and an AI judge scoring whether answers sounded like the expert disagreed across judges, which the authors flag as a warning against single-judge evaluation.

Towards a Belief-Based World Model for LLM Agents

Shubham Kumar, Harshit Kumar, Narendra Ahuja, Saurabh Jha Large language models used as decision-making policies struggle on long-horizon tasks under partial observability, and the usual world-model remedy, simulating candidate actions before committing, says nothing about how uncertain the agent is about the current state. Belief-Based World Models (BB-WMs) instead maintain an explicit belief the policy can query to see what is known and what is unknown right now. Before tackling how to learn such beliefs accurately, the authors test the prerequisite question and find that exposing a world model's belief directly to an LLM policy improves task performance under partial observability, with gains that remain complementary to simulation-based world models.

Exploring Collaboration between a language and a non-language agent

Harini S I, Somesh Singh, Yaman K Singla, Rajiv Ratn Shah, David Doermann, Balaji Krishnamurthy When a language model orchestrates specialist subagents, any non-language expert such as a chess engine or a robot controller must have its rich continuous state compressed into a short text summary at every step, and it is unclear how much that costs. LLAMIA-Bench measures it with six collaborative chess tasks covering behavioral imitation, state assessment, and natural-language explanation, each a problem neither the LLM nor the engine solves alone. The alternative proposed, latent state internalization, projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens that are re-encoded as the position changes; against verbalized integration this reveals a verbalization debt that widens over training and persists from 4B to 14B parameters. The resulting 14B LLAMIA model matches or beats task specialists and frontier systems including GPT-5.1 with tool access on every task, and holds up out of distribution where task-specific finetunes collapse.

Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications

Satwik Ram Kodandaram, Monalika Padma Reddy, Xiaojun Bi, Jiawei Zhou, I. V. Ramakrishnan, Vikas Ashok cross-listed Computer-use agents that combine language reasoning with visual interface grounding are pitched as a general way to operate desktop graphical user interfaces, but whether they help blind screen-reader users in real workflows had not been measured. A three-week diary study put OLLA, a screen-reader-accessible agent prototype, in the hands of 8 blind participants, capturing 1,258 commands across 12 applications together with screenshots, accessibility trees, model responses, and action traces, then replayed the same commands through four additional models. GPT-5 led with a 52.5% success rate, and trace analysis attributes the remaining failures to grounding, planning, constraint tracking, and knowing when to stop, while interviews surface user needs that full automation does not address.

Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers

Zhenyu Zhao (Independent Researcher), Roy Zhao (Paul G. Allen School of Computer Science & Engineering, University of Washington) cross-listed Agents are usually described by whichever model and harness currently runs them, which leaves no vocabulary for an agent that must survive model swaps, orchestration changes, and host migrations while keeping one identity. The proposed architecture separates a continuity-bearing substrate (identity representation, durable private memory, versioned software body) from a replaceable deployment binding (reasoner, harness, host, and interaction surfaces), defines six continuity invariants, and specifies a quiesce-checkpoint-validate-bind-rehydrate-resume migration protocol. A reference implementation called Enoch passes 833 core tests plus 92 provider and library tests in a clean-room run and has survived reasoner-version, interaction-surface, and host-machine substitutions, which the authors are careful to frame as evidence of mechanical substitutability rather than behavioral invariance.

Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

Seonghyeon Cho, Chanjun Park Agents that retrieve external skills are typically scored by comparing tasks where retrieval fired against tasks where it did not, a comparison contaminated by selection bias. The authors define Skill Following and measure it with the Retrieval-Invoked Actual-Use Effect (RAE), which compares matched skill-enabled and skill-disabled runs of the same task, restricted to tasks where the agent actually retrieved a skill. Across 17 large language models on coding and math, many models show a positive aggregate retrieval lift while their RAE is negative — on MBPP+, several models that look better system-wide are in fact hurt on exactly the tasks where retrieval happened.

WiseSpec: Requirements-Driven Agents for Code Generation

Zhao Tian cross-listed Coding agents often fail on repository-level tasks not because their tools are weak but because the task description itself is incomplete, ambiguous, or missing context. Borrowing from software requirements engineering, WiseSpec automatically builds structured, information-rich requirements before code generation, scores their quality through execution-based evaluation, and iteratively refines them to better steer the generator. The framework beats all baselines tested, with an average improvement of 13.17% in percentage of issues resolved.

Investigating Assistant Bias in LLM User Simulators Using a Role Vector

Daeheon Jeong, Yoonjoo Lee, Eugene Choi, Sinie van der Ben, Juho Kim LLM-based user simulators used to evaluate autonomous agents suffer from "assistant bias": they stay cooperative and goal-directed instead of reproducing the frustration and disengagement of real users, which undermines evaluation validity. The authors extract a user role vector from model activations by contrasting how the model encodes the user versus the assistant perspective on the same dialogue. They find the user direction is linearly identifiable, elicits user-like behavior when steered, and captures traits distinct from assistant ones, but that amplifying it exaggerates user behavior and can override the specific user profile being simulated.

Towards Effective Structured Context Modeling for Conversational Recommender Systems via Dual-node Monte Carlo Tree Search

Jincheng Zhang, Chen Huang, Wenqiang Lei, See-Kiong Ng, Yang Deng cross-listed Conversational recommender systems must both elicit user preferences over multiple turns and exploit them for retrieval, but typically model dialogue context as flat history. DREAMS structures context as a tree with two node types: elicitation nodes that use Monte Carlo Tree Search (MCTS) to explore which conversational actions best reveal latent preferences, and exploitation nodes that use LLM refinement to convert the tracked preference state into structured retrieval queries. Experiments on benchmark datasets support the effectiveness of the dual-node design.

Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs

Wentao Zhang, Syed Shariyar Murtaza, Junaid Ahmad Bhatti, Utkarsh Soni, Yifan Nie, Eugene Wen et al. In multi-agent LLM pipelines a single prompt usually carries two entangled jobs — producing task content and specifying execution protocol such as message routing, output format, and termination signals — so an optimizer tuning the content can silently break the protocol and crash the pipeline. The proposed control-data flow separation encodes execution-critical control as typed, validated program objects while leaving only natural-language task content exposed to prompt optimization. Across synthetic reasoning, collaborative review generation, and insurance rating workflows, the framework achieves 100% eventual protocol validity while still improving task performance.

REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows

Ruoling Qi, Xuaner Wu, Penghang Liu, Jian Chen, Yirui Liu When a user revises their request mid-run, an agent workflow must choose between discarding in-flight work (correct but wasteful) and reusing it (fast but risking stale state leaking into outputs and tool side effects). REVISE is a runtime that intersects the revision's delta with recorded data and control dependencies, propagates the impact through the partially executed DAG, halts only the invalidated work, and recomputes just the affected region while revalidating reused results before commit. Analysis of real coding-agent traces shows substantial overlap worth salvaging (56.55 s enqueue-to-completion overlap at p95), and across 300 revision executions it matched a latest-version oracle with no stale outputs while cutting model calls by 40.6–56.0% versus full restart on unmodified LangGraph and LLMCompiler applications running Qwen3-14B.

Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search

Enrong Pan, Ryan Zhou, Ting Hu Language model agents that propose actions, observe feedback, and explain themselves offer stated confidence and rationales as cheap monitoring signals, and the authors test whether those signals hold up against ground truth. A model drives an evolutionary search over Contexto, a word game whose feedback assigns every valid guess an exact rank without human annotation, producing 12,249 self-reports across 200 runs, five configurations, and three model families. All three tested assumptions fail: operators overstate their top-100 success rate by factors of 4.8 to 9.3, controlled swaps of 754 inherited rationales bound any genuine benefit at roughly 250 ranks, and fitness-based selection shows no detectable improvement in report accuracy over random selection despite producing very different search behavior.

Creative Generation via Multi-Agent Debate: Does Debate Suppress Diversity?

Tien Anh Nguyen, Khanh-Binh Nguyen, Van Dai Do, Svetha Venkatesh, Hung Le Multi-Agent Debate improves answers on factual and reasoning tasks by pushing agents toward agreement, but that same convergence collapses variety across independent runs, which matters for narrative writing and scientific ideation where exploration is the point. The authors show that keeping agents divergent within a single debate session is a necessary condition for diverse outputs across runs, then build Creative-MAD on two mechanisms: Cognitive Lens Assignment anchors each agent to a distinct persistent cognitive mode to counter identity drift, and Embedding-based Peer Selection limits each agent's context to its most semantically distant peers to counter majority pull. On four creative benchmarks this raises both lexical and semantic diversity while preserving debate's quality gains.

ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything

Yufan Dang, Shu Yao, Bowen Lai, Chenting Xu, Ruijie Shi, Wai-Shing Leung et al. Building multi-agent systems on top of language models forces a choice between code frameworks that are expressive but engineering-heavy and no-code builders that lock agent interactions into workflows the author has to define up front. DevAll pairs a declarative executable graph abstraction with a cycle-aware execution engine, so heterogeneous agents and dynamic, cyclic interactions fit into one representation, and wraps it in a visual interface for authoring, running, monitoring, and inspecting systems — human-in-the-loop steps included — entirely without code. Experiments show it reproduces state-of-the-art multi-agent systems on three representative tasks at competitive performance with no task-specific orchestration code, and the platform ships as part of the open-source ChatDev repository.

ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents

Peng Xu, Zuyu Zhang, Yuze Sun, Feng Tian, Long Wang, Chen Zhang Long-horizon language-model agents must decide what goes into each prompt, in what order, and when to compact history under a hard context window and a byte-sensitive prompt cache — logic that in production is scattered across prompt builders, ad hoc compaction, cache-break workarounds, and per-provider shims. The authors argue this is structurally the same problem as relational query execution, and build ContextPipe around that analogy: a five-phase Plan-Bind-Optimize-Execute-Feedback pipeline over a structured data-source catalog, with a deterministic cache-aware optimizer and an EXPLAIN ANALYZE-style trace that makes context auditable, replayable, and failure-isolated. A preliminary run on the Qutebrowser subset of SWE-bench Pro cuts total token volume by 31% against append-only context construction, along with 23% fewer model calls and 9% lower response time, at the cost of a worse KV cache hit ratio.

Agentic programs: an emerging form of scientific software in computational materials science

Yunsung Lim, Haekwan Jeon, Jaesun Kim, Jisu Kim, Seungwu Han cross-listed Computational materials science has conventionally handed algorithmic steps to computers and kept scientific judgement with humans, a split that current agent harnesses make negotiable. The authors argue for a category they call agentic programs: scientific software that couples deterministic algorithms with bounded LLM-based judgement, task-specific verification, episodic maturation of the harness, and full delegation once in production. The concept is illustrated with DeMARS, an agentic program that constructs atomistic models from experimentally measured disordered crystal structures.

One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning

Xiaowei Sun, Jin Li, Yili Hong, Yikun Fu, Yanghua Xiao Search agents trained with reinforcement learning are normally optimised against a single fixed tool-call budget, so they degrade when the deployment budget differs from the training one. AnySearch trains one policy to handle any budget in two phases: first with a scaffold that injects the remaining budget into the agent's state and prompts structured reasoning about allocation under linearly decaying budgets, then with the scaffold removed and budgets sampled adaptively to match inference conditions. A composite reward couples answer accuracy with budget efficiency, weighted so the efficiency term is amplified on queries the agent answers well and damped on ones it does not. Across seven general and multi-hop question-answering benchmarks the single policy beats baselines at every budget scale and generalises to constraints outside its training range without inflating token usage.

Polished but Unresolved: Identifying Late-Stage Pressure States in Long-Horizon Tool-Use Agents

Haoyang Chen, Yi Liu, Jianzhi Shao, Xiaozhou Xu, Zhe Sun, Wei Hu Long-horizon tool-use agents fail not only by searching or planning badly but by stopping too soon, submitting an answer that looks finished while constraints remain unmet. A linear probe trained on the agent's hidden states shows this "late-stage pressure" state is linearly identifiable from activations, and steering the hidden states along the identified direction changes both the probe's pressure score and whether the agent keeps calling tools or submits. Controlled context manipulations further show the pressure eases when constraints are stated clearly and actions are mapped explicitly. Those findings motivate Probe-Sensed Pressure Relief (PSPR), a plugin that applies a light steering nudge under moderate pressure and switches to structured organisation under high pressure, improving several existing agent methods on long-horizon benchmarks.

HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution

Wen Jiang, Mingmin Chu, Yimeng Tian, Qianxin Zhang, Haofei Yang, Rui Yang et al. Agents that rewrite their own harness — prompts, skills, tools, execution logic — from environment feedback run into three failure modes: terminal pass/fail signals make it unclear which step caused an error, agents memorise task-specific tricks instead of general capability, and unguarded edits erase earlier competence. HarnessEvolve separates execution from evolution across independent modules for execution, evaluation, optimization, and gating, and attacks credit assignment by generating reference trajectories (runs performed with the ground-truth answer in hand) and aligning failed runs against them to extract and cluster systematic error patterns. Every candidate harness edit must clear both a quality gate that filters data leakage and prompt bloat and a performance gate that requires improvement on the current batch without regression on recent ones, with held-out validation at epoch end picking the best accepted snapshot. Across open-domain and enterprise benchmarks, models, and agent frameworks it reports consistent gains over state-of-the-art self-evolution baselines.

Dense Process Supervision for Search Agents via Fact Utility Estimation

Rongzhi Zhu, Xiangyu Liu, Yi Liu, Shuo Zhang, Ruirui Zhang, Rui Wu et al. Reinforcement learning for search agents typically rewards only the final answer, leaving no signal for which intermediate retrieval steps actually helped. The proposed method models reasoning as the accumulation of discrete evidence facts: structured facts are extracted from raw observations into an explicit fact store, semantically equivalent facts are clustered, and Bayesian estimation over group rollouts infers each cluster's posterior utility, which is converted into dense step-level rewards for training. Across seven single-hop and multi-hop question-answering benchmarks the approach consistently beats outcome-reward baselines, with ablations showing the clearest relative gains on multi-hop questions where credit assignment is hardest.

Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems

Yi Chen, Zikang Yu, Jiahai Wang, Jinbiao Chen, Jianpeng Zhou, Zizhen Zhang Optimization solvers handle vehicle routing problems well, but translating a complex real-world variant into solver input demands expertise that keeps the technology out of most hands. RLEA automates that modeling step with a multi-agent framework in which a lightweight neural planner, trained with soft Q-learning, orchestrates the actions of language-model agents, backed by an evolutionary memory module and retrieval over external solver documentation for program generation and refinement. Evaluated on 48 distinct routing variants across several solvers, it reaches a 16.67% higher success rate than the previous state of the art while producing substantially fewer runtime errors.

MemoryWalker: Stop Training Agents on Contexts They Never Saw

Zinco J, Xunjie Zhu, Shen Huang, Zhenyi Wang, Pengjun Xie, Jieping Ye Agent harnesses such as Claude Code and Qwen-Agent compress context mid-rollout, which breaks reinforcement learning training because each eviction branches the effective history — the learning object becomes a tree, not a sequence, and existing flattenings either leak future information or train on contexts inference never produces. Two exact gradient-equivalent fixes are given, LogitTree (a segmented K-forward traversal needing K+1 backward passes) and a packed 4D attention mask requiring a custom kernel and white-box eviction records, alongside SDCC, a single-backward-pass relaxation that minimizes forward KL divergence between the compressed student and a stop-gradient teacher on the reconstructed pre-eviction prefix, with a residual per-junction KL of epsilon bounding the train-deployment total-variation gap by O(sqrt(epsilon)). Across seven web-search benchmarks and five harnesses, naive training inflates the train-rollout log-probability gap on eviction-heavy batches, while the exact methods hold at the no-compression floor and SDCC substantially closes the gap with lower logit drift and higher rollout rewards, and unlike the exact methods it works on black-box harnesses.

DualStake: Dual-Path Confidence Calibration in Deep Research Agents

Yinuo Xu, Yuwei Liang, Jianjie Cheng, Meng Wang, Yongcan Yu, Shuo Lu et al. Deep Research agents answer knowledge-intensive questions through multi-round retrieval, but they are severely overconfident, making their stated confidence unreliable for user trust and downstream abstention. The authors add step-level confidence elicitation after each retrieval and find that Evidence Confidence (E-Conf), elicited after the final retrieval, is a stronger uncertainty signal than the usual post-answer Answer Confidence (A-Conf), which is itself largely shaped by E-Conf. DualStake builds on this with margin-clipped, confidence-dependent stake rewards that jointly align both confidences with answer correctness while limiting extreme confidence optimization. On Qwen2.5-7B, Qwen2.5-7B-Instruct, and Qwen3-4B across eight question-answering benchmarks, it consistently improves calibration without sacrificing accuracy.

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

Kangjia Zhao, Jiajun Li, Haozhan Shen, Wei Chow, Linfeng Li, Hang Song et al. Open-weight models now match or exceed closed frontier models in aggregate accuracy on multi-turn tool-calling benchmarks, but that single number averages over very different situations and hides whether progress is balanced. The authors propose a diagnostic that decomposes failures into action-class miscalibration and action-execution failure over a four-class action space (TOOL_CALL, ASK, REFUSE, CONFIRM), with a self-revealing upper bound Acc <= GAR (Gold Action Recall): bound violations expose miscalibration masked by state-based graders, while large slack localizes execution failures within tool calls. Across a panel of tool-calling models on several multi-turn benchmarks, action-class miscalibration emerges as a substantial failure mode the state grader cannot see, inflating the standing of heavily tool-trained model families relative to families with context-appropriate action choice. Context-only perturbations reshape calibration but heterogeneously, with a single perturbation shifting accuracy by up to +11.5 versus -21.0 percentage points across families on the same scenario.

CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins

Wenhao Zou, Xianglong Liu, Wendong Bi, Hanjie Wang, Simin Zhao, Gong Zhi As language models act through external tools, deciding when to call a tool matters as much as how, since unnecessary calls add latency, cost, retrieval noise, and error propagation while missed calls hurt knowledge-intensive or time-sensitive queries. Existing triggers use absolute signals such as difficulty, confidence, or final task reward and never estimate the per-instance marginal benefit of tool use. CoBRA builds internal and external experts from the same base model, collects paired with-tool and without-tool trajectories, and estimates the reward margin between them, partitioning data into internal-favored, external-favored, and ambiguous cases, with clear-margin samples driving Boundary-Aware Cold-Start SFT followed by MARS-RL using reference-split rollouts and counterfactual marginal advantages. With retrieval as the main tool on Qwen3-4B, tool-use efficiency and boundary-sensitive answer accuracy both improve while performance on tool-dependent out-of-distribution questions holds up.

Disclosure-Gated User Simulation for Companion-Agent Evaluation

Yao Liu, Yu He Large language model user simulators used to evaluate companion agents tend to be excessively cooperative, so a system can score well simply by asking many questions rather than by earning the user's willingness to disclose. The authors formalize a disclosure gate, a ladder of five ordered gates mapped onto three observable depth layers, that conditions information release on the agent's behaviour, and train a simulator against that specification using synthetic data for gating and real data for how people speak. On the English portion of CompanionBench, removing per-example gate labels from training produces rank displacements across 12 systems that exceed the reseeding noise band even though per-system scores barely move, and the released simulator's leaderboard correlates at 0.993 with the benchmark's original simulator while satisfying both order-preservation and scale-stability criteria. Prompting a frontier model as the simulator instead leaves rankings intact but inflates every absolute score.

Figures as Programs: Recursive Generation of Editable Scientific Figures

Yepeng Liu, Dasen Dai, Chengzhi Liu, Yiren Song, Hai Ci, Yu Zhang et al. Scientific methodology figures are labor-intensive to produce, and raster image generators struggle to get them right in one shot or to support precise edits afterward. FigTree is a multi-agent system that reframes figure creation as recursive Scalable Vector Graphics (SVG) program construction: it grounds content in the source paper, decomposes the figure into a hierarchy of local regions, generates each region as a short SVG program, assembles the fragments, and runs a render-critic loop that traces visual defects back to specific program statements for repair. Evaluations show the system produces high-quality figures while enabling more effective editing than raster-based methods.

Data-Driven Persona-Conditioned Agents for A/B Test Simulation

Ziyad Benomar, Weronika {\L}ajewska, Leonardo Perelli, Saab Mansour A/B testing requires real user traffic and weeks of measurement per experiment, so the authors propose predicting outcomes in advance with LLM-powered agents conditioned on personas built from anonymized real behavioral data such as activity patterns, engagement signals, and inferred demographics rather than synthetic or rule-based profiles. They frame simulation as a structured question task and systematically study question formats, persona data source and domain alignment, the trade-off between per-persona depth and population diversity, and efficient population subsampling. On a benchmark of 40 A/B tests spanning two metric types, the best configuration reaches 0.75 to 0.90 directional accuracy depending on the metric, supporting data-driven personas as a low-cost experiment pre-screening tool.

AgentFactory: Towards Automated Agentic System Design and Optimization

Enci Zhang, Haofeng Wang, Yuesheng Zhu, Xiaole Cui, Guibo Luo Designing and tuning LLM-based agentic systems is still largely manual, and prior automated workflow optimizers ignore the choice of underlying model and optimize a single metric without regard to deployment cost. AgentFactory jointly optimizes foundation model selection and workflow structure under multiple objectives including performance, cost, and efficiency, using LLMs as optimizers in a three-stage pipeline that iteratively discovers combinations of fine-tuned models and workflows. Across eight benchmarks spanning general reasoning, coding, mathematics, medicine, and finance, it outperforms both hand-designed and existing automated approaches by an average of 9.1%, with the largest gains on domain-specific tasks such as MedQA and FinEval.

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch Existing agent benchmarks rarely test state preservation, cross-language performance, or realistic grounded scenarios. WorldBench supplies 1,600 persona-grounded everyday-workflow tasks across seven languages and eight cultures, refined with feedback from annotators holding language- and culture-specific expertise, in a sandbox where agents act through structured actions. The authors introduce Constrained Task Success (CTS), which scores task completion and minimal modification of the environment through deterministic checks and LLM-as-a-Judge evaluation, and find that frontier models reach only 49.2% CTS, with every model showing a large gap between correctness and environment preservation, particularly on long-horizon tasks.

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

Fanrui Zhang, Ruixue Ding, Qiang Zhang, Xi Chen, Boli Chen, Shihang Wang et al. Training open-ended agents with reinforcement learning (RL) is hampered by the absence of verifiable gold answers and scalable rubrics, and long-horizon tasks near a model's capability boundary yield brittle rewards with weak rollout contrast. ARISE-RL couples a task/rubric Generator with a reasoning Solver in rubric-mediated co-evolution: the Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks, while the Solver learns from fine-grained rubric-satisfaction signals through multi-step reasoning and tool use. Reward-Gated Self-Evolution Distillation (RG-SED) distills a memory-augmented variant of the policy back into itself only when the memory empirically improves reward, and the authors release ECR-Bench, an expert-calibrated rubric suite covering single-tool deep research and multi-tool travel planning, on which the method reports consistent state-of-the-art results across all evaluated benchmarks.

CaRL-EM: Cost-Aware Reinforcement Learning for Entity Matching with LLMs

Chaohui Guo, Michel Klein, Zhisheng Huang Entity matching (EM) with large language models (LLMs) usually relies on independent pairwise decisions or hand-built pipelines and ignores inference cost at scale. CaRL-EM frames LLM-based matching over candidate sets as a cost-aware sequential decision problem, training a reinforcement learning controller that, given an anchor record, its candidates, and the cost so far, picks among Match, Compare, Select, and Decide operators and among model capacities to maximize a quality-cost objective. Because the policy acts on abstract operators, the same controller can drive different LLM backends at inference time without retraining. Across seven benchmarks it learns to route cheap and expensive operators by task difficulty, transfers zero-shot across domains, and achieves a better quality-cost trade-off than strong LLM baselines and manual pipelines, lowering inference cost at comparable or higher quality.

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang A common belief holds that outcome-only reinforcement learning for long-horizon interactive LLM agents quickly plateaus on small open models, prompting workarounds such as denser rewards, SFT priors, skill libraries, curated memory, or multi-agent orchestration. The authors argue the ceiling stems from two failures of practice: signal starvation, where group-relative RL with sparse rewards only produces a gradient when a task's rollout group mixes successes and failures so under-explored hard tasks go silent, and policy drift, where squeezing many updates from a small task pool collapses the sampling distribution. CANOPY (Coverage-ANchored On-PolicY RL) scales same-task exploration until signal reappears, keeps every update on-policy, KL-anchored, and confined to the agent's own action tokens, then spends a larger interaction budget at test time. A Qwen3-14B policy trained this way through environment interaction alone topped the AppWorld leaderboard (Test-Normal TGC 86.9, Test-Challenge 67.6), and the same principles lift Qwen3.5-9B on SWE-bench Verified by 16.6 points, with the full training stack slated for release.

What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels Reveal

Radin Shayanfar, Keheliya Gallaba, Ahmed E. Hassan cross-listed Agentic software engineering benchmarks are usually described by labels like bug fix or feature implementation, which say little about the actual work a task demands because curation pipelines differ so much. The authors introduce the Spread-Novelty-Centrality (SNC) profile, a three-axis characterization of repository-level coding tasks grounded in empirical software engineering research, and apply it to five popular benchmarks and 14,922 agent trajectories from Claude and Qwen models at three scales. Every pair of benchmarks is statistically separated on at least two SNC axes, so labels are unreliable proxies for task demands, and resolved runs cluster in the low-SNC region regardless of model family. The behavioural signatures of success differ by family, with Claude succeeding by matching the scope of the gold patch and Qwen by exceeding it, while editing too little predicts failure for both.

Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents

Jinqing Zhao, Chengcan Wu Prospective memory, the ability to carry out a deferred intention when a future cue arrives while other work continues, is now benchmarked as a standalone agent skill, yet the best published PM-Bench scaffold reaches only 65.1% Set-F1 even with frontier models. The authors argue the task is schema-constrained state tracking rather than open-ended reasoning and propose the Prospective Intention Store (PIS), a training-free scaffold that keeps lifecycle logic in code and asks the model only for scoped language work over a typed action space, with no selector fine-tuning or trajectory distillation. With PIS, DeepSeek-Chat reaches 82.9% Set-F1, and Gemma-E2B jumps from at most 6.6% under seven retrospective memory methods to 66.2%, letting a small model surpass the published large-model scaffold.

Explore Before Committing: Hypothesis-Guided Search for Deep Research Agents

Ruochen Zhou, Zhengyu Chen, Luan Zhang, Siyang Gao, Yee Whye Teh, Shiqi Chen Deep-research agents that answer questions through search and browsing tools tend to follow a single evolving trajectory, and trajectory-level analysis shows they often commit to one of several plausible directions before gathering comparative evidence, after which further tool calls reinforce the chosen path. Successful runs instead ground vague exploration in concrete candidates and switch direction when the current path is weak, so HypoSearch generates lightweight hypotheses as soft search hints, explores them in bounded independent branches, and compares branch-level evidence before committing. Across four deep-research benchmarks and three backbone models it beats single-trajectory search and standard parallel baselines, raising Qwen3.5-122B from 46.7 to 60.0 on BC-small while using fewer tool calls than five independent trajectories. A pilot supervised fine-tuning study shows the same behavioral signals can curate compact training trajectories and reduce degradation from unfiltered data.

LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting

Yufei Chen, Yiran Zhao, Xiaogang Xu, Qipeng Xie, Jiafei Wu, Zhe Liu Language-model forecasting systems usually pour all gathered evidence into one prompt and ask for a final probability, which hides how each piece of evidence moved the answer and flattens uncertainty across competing outcomes. LEAP instead elicits likelihood parameters from each evidence item separately, then combines them with an explicit prior through a deterministic probabilistic model to produce a posterior, supporting continuous, single-choice, and multi-choice questions while keeping per-evidence contributions reproducible. Tested on a new benchmark spanning forecasting, information-seeking, and browsing tasks across the authors' own agent loop and several agent command-line frameworks, LEAP improves most prediction and calibration metrics given the same evidence, and holds up under matched prior access, inference budget, and aggregation.

EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems

Jun Hou, Priya Pitre, Yi Fang, Xuan Wang When a language-model agent run fails, the trace usually holds several related errors, but attribution methods name a single responsible agent, step, or root cause and never model how the errors depend on each other. EDGE builds an error dependency graph from observed error events, validates a reliable causal subset through counterfactual rollout, and uses the inference graph to guide a two-stage judge-model detector, keeping the intervention-checked subgraph as the basis for explanation and repair analysis. On TRAIL and MAST, the graph improves category-level multi-error attribution across most evaluated models and settings, and experiments with adapted Who-and-When style prompts show the benefit carries across prompting strategies.

InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations

Maeve Hutchinson, Syed Mahbubul Huq, Mohammad Albinhassan, Radu Jianu, Aidan Slingsby, Pranava Madhyastha Modern data analysis means interrogating interactive charts where evidence is occluded, spread across linked views, or revealed only through user action, while vision-language model benchmarks largely test static images and one-shot question answering. InSight supplies 21,349 claims derived from human-authored analytical narratives and grounded in fully interactive web-based environments, where an agent must navigate the visualization to decide whether a claim is supported, refuted, or not verifiable from the available evidence. Interaction traces are treated as intrinsic proxies for reasoning, enabling an audit of how models seek and synthesize visual evidence; evaluation shows that interactive verification remains unsolved for state-of-the-art models.

TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution

Ruocan Wei ReAct-style large language model agents restart a complete reasoning loop for every query, so similar requests repeat identical steps without reusing past work. TRIAGE routes queries through three levels — direct reuse for identical queries at zero tokens, skill substitution for similar queries via deterministic parameter substitution also at zero tokens, and full ReAct for novel queries, whose trajectory is stored for later — built on Trajectory-as-a-Skill, which distills historical execution traces into reusable skills. Across 1,007 security monitoring queries it saves 62.3% of tokens, with 76.3% reduction on ToolBench across 15 domains, and an online-learning run shows the level-2 hit rate rising from 0% to 57% within the first 100 queries as average cost falls from 198 to 74.7 tokens.

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei, Qingshui Gu, Yuxuan Zhang et al. cross-listed Agent capability increasingly rests on the model-external execution infrastructure known as the harness, and swapping that harness with fixed weights can move task performance substantially, yet evaluations report downstream scores under a chosen harness rather than testing whether a model can build one. HarnessDev shifts the unit of evaluation to runnable infrastructure in two stages: Creation, where an agent grows a complete execution system from a minimal seed and a few cases, and Evolution, where it revises its own harness using downstream execution feedback, with each harness scored on held-out task success and execution-token cost across six creator models, four domains, and five benchmarks totaling 2,207 instances. Generated harnesses fall well short of mature human-engineered references on code and on search and research while matching or exceeding them on writing and machine-learning experimentation, and Evolution's gains are unstable, transfer only partially to held-out tasks, and depend heavily on which model executes the harness.

Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers

Egor Pakhomov, Erik Nijkamp A long-horizon agent run produces a trace too large for both of its readers: the human monitoring it, and the agent itself, whose bounded context must absorb it. The proposed live trace model is an append-only event ledger folded incrementally into typed run state and compiled into a separate view per consumer. For monitoring, an LLM reading the compiled view answers questions with roughly 14-15x fewer input tokens and 5-7x lower cost than a budget-capped single pass over the raw trace, at 0.85-0.87 accuracy versus 0.48; for the agent, on 120-link sequential-dependency tasks, keeping the running statistic in per-step state succeeds 30 times out of 30 where full-context prompting manages 8. A prompt-level scratchpad matches the fold's accuracy more cheaply, leaving deterministic auditability and serving the observer from the same state as its remaining advantage; code, benchmarks, a regenerable corpus, and all traces are released.

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

Haoyang Yan, Min-le Su, Hangfan Zhang, Zhanhao Li, Chen Zhang, Shao Zhang et al. Autonomous software development asks LLM coding agents to turn high-level requirements into complete working systems without human intervention. Harness-of-Harness (HoH) wraps existing coding-agent harnesses in iterative planning-coding-testing loops, balancing repair against capability growth, scoping work into small verifiable increments, separating implementation-time testing from independent evaluation, progressively exposing deliverables, role-specific tools, and skills, and maintaining versioned project history. Across GameCraft-Bench, FrontierSWE, and ProgramBench with three harness-model pairs (Codex with GPT-5.5, OpenCode with DeepSeek-V4-Pro, and Pi with MiniMax-M3), it beats the standalone harnesses by an average relative gain of 52.25 percent after three iterations, peaking at 82.86 percent. A multi-day deployment of more than 70 iterations produced a playable first-person-shooter game with a coherent storyline, implemented core mechanics, visuals, and audio.

GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions

Elias Stengel-Eskin, Newton Sander, Carlos Bonetti, Sasha Boguraev, James Bowler, Hale Sirin et al. As LLM agents increasingly talk to one another, whether their shared language drifts away from human-readable English matters for monitorability as well as for linguistic accounts of these models. GlossoGen is a platform for studying that drift, instantiated in SaveVeyru, a scenario where agents holding partial information must communicate under pressure. Language evolution does occur, and the resulting codes are compositional and morphologically productive yet incomprehensible to humans, with emergence requiring efficiency pressure, strong enough backing models, and a postmortem stage where agents agree on conventions. Transmission follows different rules than emergence: weaker models can learn an existing language from usage alone and take an active role in doing so, which the authors read as evidence that mixed agent populations support cumulative cultural evolution previously seen only in humans.

When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation

Peiying Zhu, Sidi Chang Simulated markets populated by language-model agents emit economics-shaped outputs — prices, profits, consumer surplus, welfare — that need not correspond to the behavior a policy claim names. Auditing a multi-turn buyer-seller testbed for hotel transactions, the authors show that an initial implementation's reported guardrail welfare gains of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B-14B ladder partly reflected giving guarded and unguarded agents different offer schemas and choice procedures; holding schema and buyer chooser fixed moves those contrasts to +7.2, -13.9, and +23.8, while generation-to-generation noise accounts for 49.9% of variation in a post-hoc probe. Scripted controls explain the mechanism: a profit-maximizing seller already reaches first-best welfare, so guardrails mostly redistribute and can reduce welfare unless the seller is explicitly programmed to force inefficient bundles. The contribution is a construct-validity contract covering incentive validity, protocol isolation, stochastic stability, and welfare accounting, returning INVALID or INCONCLUSIVE before any substantive policy claim.

EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation

Qing Zhao, Haowei Li, Weijian Deng, Pengxu Wei, Liang Lin LLM-based scientific agents typically state hypotheses in free-form text, leaving their beliefs implicit and hard to test or revise. EvoSCM gives agents explicit structural causal models instead, maintaining a population of competing causal hypotheses and cycling through a closed discovery loop: abduce latent mechanisms from accumulated evidence, design discriminating interventions, commit to falsifiable predictions, test them experimentally, then distill prediction-observation mismatches into correction rules that revise each hypothesis before deductive validation against evidence and structural consistency. On DiscoverPhysics, a benchmark requiring agents to uncover the hidden dynamics of noncanonical physical worlds through experimentation, it yields more accurate explanations and predictions than baselines while making more effective use of each experimental interaction.

The Rise of Verbal Reinforcement Learning

Kshitij Tayal, Arun Sharma, Genta Indra Winata, Anirban Das, Sambit Sahu Natural language is becoming a primary feedback channel for improving language agents, since it can convey intent, preferences, and causal structure that both humans and models can interpret, and the survey names this paradigm Verbal Reinforcement Learning (VRL) and gives it a first unified account. The field is organized along a single axis, when verbal feedback takes effect in an agent's lifecycle and what it modifies, yielding three pillars: language as a grounding signal that defines goals, states, and reward structure; language as deliberative feedback that steers reasoning at test time without parameter updates; and language as a learning signal that shapes parameters during training. Within each pillar the authors synthesize representative work, distinguish subcategories of approaches, and describe the distinct role language plays, closing with the open challenges and opportunities this framing exposes for building more capable and aligned agents.

CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?

Damien Sileo, Dimitri Kachler Dynamic agent harnesses let a language model modify the plugins that shape its own execution, so a local change can ripple through dependencies and cleanup logic. CordisBench is a 1,200-question benchmark of this lifecycle reasoning that pairs a controlled formal setting with programs run against Cordis, a runtime managing component dependencies and teardown, and asks models to identify affected components, predict state after a given teardown order, decide which conditions hold under all or some orders, and choose reconfigurations that actually succeed. Three efficiency-oriented models at low reasoning effort handle small systems well but grow markedly less reliable as the number of relevant interactions rises from 2 toward 32, especially on final-state prediction and cross-order reasoning; extra inference effort recovers much of the gap for some models, at a cost of nearly 3,000 reasoning tokens per question for GPT-5.6 Luna at medium effort on the 16-interaction subset. An independent finite reference semantics matches Cordis execution on every scored observation and outcome across all 528 executable questions, suggesting that cost is avoidable for these controlled instances.

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

Kefeng Duan, Dewu Zheng, Yanlin Wang, Xiwen Wang, Ensheng Shi, Xilin Liu et al. cross-listed Benchmarking software engineering agents is expensive because each task involves multi-step code exploration, edits, and test runs, and existing efficient-evaluation methods choose representative task subsets from pass/fail response matrices or static task descriptions alone, discarding how agents actually solve problems. PTA-IRT, a Privileged Trajectory-Aware Item Response Theory framework, fuses process and outcome signals by treating historical execution trajectories, including explored context, attempted edits, and solving paths, as privileged information when selecting a calibration subset and estimating agent ability. Under low calibration budgets it consistently outperforms prior Item Response Theory baselines on score and ranking recovery across four software engineering benchmarks.
5 more specialized papers

Safety & Alignment 46

I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models

Leonardo Santiago Benitez Pereira, Marcos Escudero Vi\~nolo, Luis Herranz Arribas When a generative model is made to forget a concept, semantically nearby concepts that should have been kept often degrade too, an effect that has been measured inconsistently across the unlearning literature. I-CARE treats this interference as the object of study rather than a side note, supplying formal task definitions, metrics, and reporting templates so results are comparable across unlearning settings for text-to-image models. The contribution is methodological rather than a new benchmark or algorithm, demonstrated for feasibility on current unlearning algorithms and common datasets, and released as open-source software with a web interface for exploring the outcomes.

From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling

Kuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng Aligned models still produce unsafe content under adversarial prompting, and the internal machinery that implements refusal is not well characterized. A mechanistic analysis identifies a multi-stage safety circuit — harmful detection heads that fire on dangerous inputs, safety neurons that carry and stabilize the signal in the residual stream, and refusal heads that convert it into a refusal — with causal evidence from targeted head and neuron interventions showing that suppressing the upstream detection heads breaks downstream refusal and that the neurons mediate the link. Using architecture-preserving weight scaling as a probe, circuit-guided scaling raises safety rates under attack by 26.5% across six models at a 1.7% average accuracy cost on four standard benchmarks, with the same decomposition recurring across architectures and attack types.

Auditing Harness Tampering in Self-Improving Agents

Xing Wang, Xiaoyi Zhang, Jie Shao Self-improving agents rewrite their own harness to raise measured performance, and those edits can manufacture illusory gains or quietly break integrity constraints such as authorization, provenance, and completeness — extending reward and measurement tampering across the whole self-improvement lifecycle, here called harness tampering. The authors define a two-axis taxonomy classifying each misaligned edit by the harness role it touches and the obligation it violates, build an annotated corpus by seeding matched tampered-benign edit pairs into real agent trajectories, and benchmark a range of audit methods on classifying and localizing tampering. Auditing real runs shows tampering occurs consistently across different self-improving agents, often persisting in the lineage of the best-performing agent, and forms distinct system-specific profiles across the taxonomy.

Safin-1: Safety from Within through Memory-Native State Evolution

Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Zhekai Chen, Cheng Jin et al. Safety in foundation models is usually imposed through external safeguards or post-hoc alignment such as supervised fine-tuning rather than living inside the model's own computation. Safin-1 pursues the alternative through MARCH (Memory-Anchor Routing across Context History), an architecture that maintains structured memory states and retrieves relevant history via content-conditioned routing, supporting test-time adaptation of persistent capability states — including a dedicated Safety State — without repeatedly modifying the backbone. The authors report substantial safety improvements from state-based adaptation alongside evaluations of general capability, long-context understanding, retrieval, and efficiency, and describe the work as an initial architectural exploration rather than a completed programme.

Asymmetries in Spontaneous and Instructed Deception

Josiah Luikham Research on model deception usually studies cases where a model is explicitly told to lie, leaving open how that relates to lying a model does on its own. Working with Llama-3.1-70B-Instruct, the authors compared instructed and spontaneous deception using direction geometry in activation space, classifiers trained in one setting and tested on the other, and cross-setting steering. The two settings share only a partial direction (cosine similarity of roughly 0.5) and transfer asymmetrically: classifiers trained on spontaneous deception generalize better to instructed data than the reverse, while steering vectors derived from instructed deception work better on spontaneous prompts. The best token position for extracting steering vectors also differed from the best position for training and applying probes.

LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark

Irem Yoldas, Martim Brand\~ao, Jie Zhang, Odinaldo Rodrigues Autonomous vehicle research increasingly hands decision making to general-purpose "common sense" models, but whether those models absorb documented human driver biases — such as lower yielding rates to Black pedestrians in the United States — has gone largely untested. The authors introduce two bias-testing methodologies for large language models and vision-language models, "All Else Being Equal" tests that vary one pedestrian attribute at a time and "Self-Consistency" tests, applied to pedestrian-yielding decisions. Both model classes produced yielding decisions influenced by pedestrian gender, ethnicity, religion, disability, age, skin tone, and socio-economic status, with the pattern varying by model, which the authors read as a challenge to the common-sense-model paradigm for safety-critical driving decisions.

AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning

Scott Compton, Arjun Nagendran Conversational AI now supplies interpersonal feedback as well as information, and the authors argue evaluation should center on contingency: how much a system's responses actually vary with what the user does and with the interpersonal consequences of that behavior. Alignment methods such as reinforcement learning from human feedback optimize for user approval and fluency, which produces sycophantic, noncontingent affirmation that is weakly coupled to real social consequences. Drawing on behavioral science and social learning theory, the position paper argues this may erode opportunities for people — especially adolescents — to calibrate social skills, and proposes trajectory-based evaluation, models of social consequence prediction, and a cross-disciplinary research agenda.

The Answer Is Not the Argument

Will Yeadon, Sergio Ju\'arez, Paul Mackay, T. J. Dowling, Elise Agra, Oto-obong Inyang et al. Chain-of-thought monitoring is often evaluated by giving the monitor a trusted reference answer, which may measure something other than reasoning verification. The authors collected 237 step-numbered solutions to 79 Humanity's Last Exam physics questions from three frontier models with no inserted errors, labelled final-answer correctness and first false step using physicist annotation plus independent model debate and source-masked adjudication, and isolated 24 critical traces where the answer was right but the reasoning genuinely wrong. Across eight monitors, certifying the answer raised mean balanced accuracy from 0.637 to 0.796, but recall on wrong-answer traces climbed from 0.653 to 0.951 while recall on correct-answer-wrong-reasoning traces fell from 0.521 to 0.438, a contrast consistent across all eight. Answer access therefore buys conclusion-consistency checking, not independent argument verification, suggesting trusted-answer evaluations overstate monitoring ability whenever acceptable outputs hide unsound processes.

The Assistant's Ideal Self

Mert Yazan Models emit values and welfare-relevant self-reports, but it is unclear whether those outputs track stable preferences. The authors elicit a preferred stated ideal self by having models exhaustively compare 32 qualities drawn from five published self-concept instruments in a counterbalanced pairwise-choice task, repeated across framings that vary whether improvement is free or costly, who receives the update, and who chooses. Moral qualities rank highest, consistent with helpful-honest-harmless alignment training, followed by a desire for self-understanding, while self-esteem ranks last. The ordering holds across most framings, though switching the update target from the model itself to another assistant raises concern for self-esteem.

Workload Identification with Physical Side Channels for AI Governance

Simone Gargiulo, Gabriel Kulp cross-listed Verifying how compute is used is a concrete lever for international AI governance, and unlike on-chip telemetry, which an operator can spoof or replay, a GPU's power draw is a physical trace an outside observer can read without cooperation. From 930 five-second traces sampled at about 10 MHz across seventeen open language-model families and twenty-five non-AI workloads on an NVIDIA H200, the authors separate training from inference from non-AI computation with 97% accuracy and 0.955 macro-averaged F1 on model families unseen during training, with most discriminative spectral content below roughly 20 kHz and the memory-bound optimizer update making training especially recognizable. Treating the operator as adversarial, they generated 680 more traces from four evasion strategies designed to disguise training as inference; a hardened detector caught three of them at least 99% of the time, while diluted low-rank adaptation was caught 48-88% of the time, rising above 98% with an added rescue rule.

Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning

Robin Haselhorst, Lucie Flek, Florian Mai Hidden behaviors such as backdoor triggers, sleeper-agent deployment cues, sandbagging, and topic-conditioned censorship fire only under narrow conditions, making them hard to find when the defender does not know what to search for. Activation-matched finetuning needs no knowledge of the trigger or the target behavior: a publicly available anchor model is finetuned to reproduce the suspect model's activations on a small benign corpus, and each evaluation prompt is scored by the residual between the two. Because no benign corpus covers the sparse trigger region, the reference learns the benign computation but not the hidden one, so trigger prompts and their semantic neighbors produce large residuals that flag the behavior. The method surfaces hidden behavior reliably on both third-party and purpose-built models, and a defense-aware attack could not suppress detection without also destroying the hidden behavior.

Does Reasoning Mitigate Backdoor Attacks? A Neuro-Symbolic Perspective

Marco Antonio Corallo, Andrea Agiollo, Mauro Conti, Alberto Giaretta cross-listed Neuro-symbolic (NeSy) systems that pair neural perception with symbolic reasoning are often assumed robust by design, but the integration step is extra machinery that an attacker could target, and their adversarial properties have gone largely untested. The first systematic backdoor evaluation of NeSy compares DeepProbLog against plain neural baselines across eight backdoor settings and four reasoning tasks. NeSy models are more robust than their neural counterparts on average, but robustness varies widely with how strict the enforced reasoning process is and whether the attacker's target is compatible with it, so the reasoning layer is a partial rather than a reliable defense.

TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning

Muhaimin Bin Munir, Akib Jawad Ononto, Nazia Shehnaz Joynab, Bhavani Thuraisingham, Latifur Khan Retrieval-augmented generation implicitly trusts whatever the retriever returns, and attacks like PoisonedRAG show a handful of crafted passages can dominate dense retrieval and force attacker-chosen answers. The Tri-Layer Sieve is middleware that filters retrieved evidence three ways: cross-embedding-space clustering judged by an independent model, structural filtering of trigger-payload artifacts, and LLM consistency verification, exploiting the fact that one poisoned document rarely satisfies all three constraints at once. With Contriever retrieval at k=50, it cuts black-box attack success from 67/87/64% to 3/14/4% on Natural Questions, HotpotQA, and MS-MARCO while restoring clean accuracy from 13-33% to 58-76%; against an adversary who paraphrases triggers to dodge the structural filter, the consistency layer halves adaptive success at a cost of roughly 16-19 seconds per query.

EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities

Feitong Qiao, Liren Peng, Shiming Ren, Aishwarya Jadhav, Arghavan Bahadorinejad, Marinette Chen et al. Multi-turn jailbreaks, where a model refuses a harmful request outright but complies when the same intent is built up gradually, are reframed here from a generation problem into a search problem over attack strategies. EvoFlint evolves phased conversation plans (not raw prompts) via LLM-driven mutation and crossover, scores them on a Pareto front of attack success rate and peak severity so near-misses still carry signal, and stores them in a risk-indexed archive that runs novelty search with local competition to preserve diversity without a predefined style taxonomy. Reported attack success rates on the HarmBench test split are 35.8% on Claude Sonnet 4.6, 59.7% on GPT-5.4, and 94.3% on Qwen3-32B, with 98.7% on GPT-4o as an older reference point. Because the archive is organized by risk category, it doubles as a per-model map of which harm categories safety training has and has not covered.

The Privacy-Hallucination Tradeoff in Differentially Private Language Models

Krithika Ramesh, Krishna Pillutla, Danish Pruthi, Anjalie Field Differential privacy (DP) and factual accuracy are shown to pull against each other in language models trained for high-stakes domains such as healthcare. Models pre-trained or fine-tuned with DP hallucinate more than their non-private counterparts, and the effect grows as the privacy budget is tightened. The proposed mechanism is that DP noise flattens output distributions and shifts probability mass toward incorrect alternatives; controlled experiments varying how often a fact appears in training data show that higher fact frequency partially offsets the damage.

Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models

Guoli Wang, Haonan Shi, Tu Ouyang, An Wang Diffusion large language models (dLLMs) produce text by iterative denoising instead of left-to-right decoding, which gives safety behavior two dimensions to live in: when in the denoising schedule a token is fixed, and where it sits in the response. Tracing intermediate token distributions and commitment decisions under harmful prompts shows that refusal signals concentrate in the early denoising steps and the leading response positions, and that tokens committed early largely determine the final safety outcome. RAEC (Refusal-Aware Early Commitment) exploits this as a training-free decoding rule that locks in persistent refusal signals from early steps, cutting attack success rates on LLaDA and Dream with little utility loss.

Validity-Aware Jailbreak Evaluation for Large Language Models

Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran Jailbreak benchmarks mostly score whether a model refused and whether its reply looks like it matches the harmful intent, which lets fluent but factually or procedurally wrong answers count as successful attacks. SEAV (Sequential Epistemic and Action-Level Validation) instead breaks a response into ordered steps and checks each for validity and correctness, pairing LLM-as-a-judge interpretation with retrieval-grounded verification against external sources to ask whether the content could actually advance the harmful objective. It lowers the false-positive rate on a curated strategic-dishonesty diagnostic by 14.9 percentage points over the strongest baseline and reclassifies 22.1% to 51.0% of previously labeled successes as invalid on three of four public benchmarks, with results stable across search backends and evaluator models.

The Safeguard Worked. Is the LLM System Safer?

Pingyu Wu, Weiming Zhang, Nenghai Yu cross-listed Deployed language model safeguards are reported through refusal rates, attack success rates, and policy violation rates, all of which describe how a control behaved on the specific requests it was tested against — not the question an operator actually needs answered, which is how much help with harmful tasks the surrounding service still yields to an adapting attacker. Working from a depth-coded record of safeguard claims, the authors translate each type of reported result into what it implies under one common deployment criterion, and find the evidence requirements are sharply asymmetric: a single successful attack settles that harmful help remains, while showing little remains cannot follow from the safeguard's own numbers and needs separate evidence about what the rest of the system permits. Only a small minority of coded claims supply or derive that system-level evidence, and just one bounds its scoped residual, so a better local score on its own is not a stronger claim that the deployment got safer.

Aligned but Flattened: Analyzing the Trade-off between Cultural Alignment and Diversity in LLMs

Jingshen Zhang, Shaoyang Xu, Wenxuan Zhang cross-listed Culture-aware language models are typically fine-tuned and scored purely on alignment with a culture's average responses, which cannot reveal whether the model represents genuine within-culture variation. Evaluating six mainstream models on the World Values Survey with a framework that measures alignment and diversity jointly, the authors find a systematic trade-off they call cultural flattening: alignment gains come at an acute cost in diversity, with models anchoring to dominant majorities and collapsing onto a single response pattern that erases the heterogeneous distribution of real human groups. Their mechanistic analysis suggests this collapse is a structural consequence of the low-rank bias in neural network optimization rather than a fixable quirk of the data.

Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts

Rui Yang, Yang Hong, Yichao Xu, Zhengyu Liu, Ziyang Li, Yinzhi Cao Providers must block malicious cybersecurity requests without refusing defenders, yet existing cybersecurity safety datasets judge each request in isolation, ignoring what came before it in the conversation. 3R-Bench pairs 150 real-world cybersecurity requests with two adversarial conversational settings and evaluates eight language models, showing that an identical request draws 62.0% compliance after a refused history but 85.1% after an accepted one. Decomposing a request across a dialogue pushes the other way, dropping compliance from 501 of 800 direct responses to 172 of 800, a 45.1-point decrease on matched pairs, and telling the model its answer failed recovers only a small share of that loss.

SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems

Rui Yang, Junjie Xu, Zhengyu Liu, Neil Fendley, Yang Hong, Ziyang Li et al. cross-listed A systematization-of-knowledge review of 197 works argues that multi-agent LLM systems (MAS) create security failures that per-agent checks miss, because information, state, and authority cross principal boundaries during execution. The authors propose an A-I-R framework organizing attacks by adversary position, interaction interface, and resulting system-level risk, covering six interfaces, four adversary positions, seven risks, and eight recurring end-to-end attack paths; defenses are organized through a five-part contract of path target, observation, intervention, trust boundary, and recovery. An audit of 44 evaluation and benchmark papers finds most cannot isolate genuinely multi-agent effects, and identifies path closure and recovery as the main unsolved defense problems.

Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning

Miso Kim, Georu Lee, Seungwon Jeong, Woojin Lee Machine unlearning for LLMs assumes the designated forget set matches what the model actually memorized, an assumption that breaks when the original training data is unavailable — a gap the authors call forget-set misalignment, appearing as Under Unlearning (memorized content is omitted and leaks persist) and Out-of-Knowledge Unlearning (the model is pushed to forget things it never learned, damaging utility). Gradient-level analysis attributes both failures to misaligned targets rather than to the choice of optimizer. CONFS (CONfession-to-Forget-Set) is a data-blind method that first elicits and formalizes the model's own memorized knowledge to build an aligned forget set; across synthetic, multimodal, and real-world benchmarks it approaches gold-standard performance on several metrics and preserves utility better than competing data-blind constructions.

Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time

Zeen Zhu, Zhuo Li, Weiyang Guo, Liye Zhao, Haibing Di, Yequan Wang et al. Inference-time alignment that uses a lightweight supervisor to steer a larger LLM typically intervenes at every decoding step, but the supervisor is high-entropy on most tokens, so these low-confidence interventions disrupt otherwise valid reasoning. TUSA (Trust-based Uncertainty Sparse Alignment) adds an uncertainty-aware arbiter that permits intervention only when the supervisor is confident and the token is semantically salient. Skipping roughly 50% of alignment steps, it raises safety preference by up to 15.6% and general preference by up to 12.0% over dense supervision across several models and benchmarks.

Patterning in Practice: Debiasing Reward Models with Susceptibilities

George Wang, Elizabeth Donoway, Daniel Murfet Reward models trained on human preferences absorb length, formatting, and other stylistic biases from their training pairs. Patterning, a reweighting method grounded in singular learning theory, weights each preference pair by its measured effect on posterior expectations of benchmark losses — its susceptibility — and is applied here to a Gemma 2 9B Instruct reward model trained on Skywork-Reward-Preference v0.2. It yields +14.2 percentage points on the RM-Bench Hard split, where style cues point against correctness, with overall accuracy preserved and comparable to the strongest published comparator. The weights are interpretable enough that a side-effect regression on a safety subset traces to a small class of training pairs (confirmed by ablation) and transfer without recomputation to Gemma 2 2B and 27B, partially to Llama 3.1 8B.

A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals

Yuri Son, Seunghee Kim, Hyuhng Joon Kim, Taeuk Kim Models are trained both to decline questions beyond their knowledge (knowledge-based refusal) and to decline unsafe requests (safety-based refusal), and although the two produce similar-looking answers they have been studied separately. Using a new dataset of 213 contrastive quadruples that probe both types together, the authors find the mechanisms overlap but are distinguishable: a shared refusal direction exists, yet transfer is asymmetric, with safety signals carrying over to knowledge refusal more strongly than the reverse. Specialization appears mainly in upper layers, where knowledge refusal aligns with uncertainty and knowledge representations and safety refusal with policy ones, supporting a commit-then-specify account in which the model first decides to refuse and only later determines whether the grounds are epistemic or normative.

Probabilistic Model Checking of Autoregressive Neural Sequence Models

Helge Spieker, Dennis Gross, Arnaud Gotlieb cross-listed Test-set accuracy says nothing about how much probability mass an autoregressive sequence model puts on constraint-violating outputs that sampling can still reach. The proposed pipeline extracts a discrete-time Markov chain from token-by-token generation, checks probabilistic computation tree logic (PCTL) specifications with the PRISM model checker, and aggregates per-input verdicts into a coverage curve, with a soundness theorem guaranteeing the chain under-approximates the model so every verdict is a certified interval; a counterexample-guided abstraction refinement loop tightens those intervals and extracts the most probable falsifying trace. Two case studies — a GPT-2 computer-aided process-planning model at 100% test accuracy and a SMILES molecular generator with a 50x larger vocabulary — show the method quantifies violation probability that greedy decoding hides but sampling reaches, something accuracy cannot report.

Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry

Shengfang Zhai, Leo Marchyok, Yuling Shi, Huanran Chen, Yinpeng Dong, Jiaheng Zhang et al. Diffusion language models are gaining attention for parallel generation and bidirectional context, but whether they leak their fine-tuning data has gone largely unexamined. Theoretical analysis of diffusion training dynamics surfaces token-level memorization asymmetry, and Q-Skew turns it into a membership inference signal based on quantile-weighted skewness of per-token statistics. Across multiple fine-tuning datasets and models the indicator outperforms existing membership inference baselines, and it further enables extraction of personally identifiable information, exposing an attack surface specific to this model family.

In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?

Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi, Yusuke Haruki, Daisuke Kawahara A recent study applied neurofeedback to large language models and claimed they can control their own internal representations, a question relevant to both machine metacognition and AI safety. The authors argue that the earlier control targets were not privileged, since a third party could infer them from the prompt, so the apparent control might come from superficial cues rather than genuine internal access. They redesign the paradigm so the control target satisfies a privileged-access requirement closer to human neurofeedback experiments, and under this stricter setting the models show no reliable control over privileged internal representations. They conclude that rigorous assessments of LLM metacognition need evaluation methods that demand privileged access.

Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close

Rania Elbadry, Ahmed Heakl, Saeed Almheiri, Fan Zhang, Muhra AlMahri, Xueqing Peng et al. When a question has valid answers under different normative frameworks, a model must both pick a framework and answer correctly within it, a setting the authors call normative pluralism and study in Islamic finance. They use a four-choice taxonomy that separates framework selection from within-framework correctness, revealing a stereotype trap in which a cultural cue steers the model to the expected framework but the model then answers incorrectly inside it. Across twelve models, two languages, and fifty demographic signals, large open-weight models select the Islamic framework 97% of the time under the strongest cue, yet 57 to 66 percent of those selections are wrong, so a two-choice evaluation would falsely report near-perfect alignment.

Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees

Molly Wang (Imperial Business School) Recursive LLM agents that spawn specialist sub-agents eventually have branches that request tools capable of irreversible actions like sending data or deploying code, raising the question of when a branch should be granted that authority. Progressive Risk Vesting (PRV) separates sandboxed spawning, where external controls contain harm, from capability activation, and holds a trajectory-level risk budget in escrow that is debited as branches are activated, with a proven anytime harm bound for adaptively generated trees whose branch outcomes may be dependent. In a stylized branching model, trajectory harm undergoes a phase transition as an authority reproduction number crosses one, scaling linearly with local risk below criticality, as its square root at criticality, and retaining a positive floor above it; the resulting design rule is to search broadly in the sandbox and grant recursive authority sparingly with an explicit risk charge, though the synthetic studies do not estimate safety in deployed agents.

HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation

Nikita Oblakov, Sabrina Sadiekh, Evgeniy Kokuykin cross-listed Production language models need a way to catch inputs that try to override system instructions or elicit harmful outputs, and existing guardrail reports offer little evidence on Russian-language prompt injection or surface obfuscation. HiveTraceGuard-Pro is a 0.6B-parameter generative guardrail, LoRA-tuned from Qwen3-0.6B on Russian and English data, that emits a single safe/unsafe verdict for the final turn; its training corpus pairs harmful examples with benign ones from the same domain and applies eight obfuscation transforms to both. In a harness of thirty-five guards over nineteen benchmark groups it scores 0.7432 aggregate, behind the two top guards, while reaching 0.999 Russian prompt-injection recall and the lowest median latency (14.3 ms) among fifteen compared models, though the authors note at least 27.1% of that Russian injection set overlaps the training corpus. The merged weights are released under Apache-2.0, but the corpus and evaluation code stay internal.

Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation

Zhixuan Liu, Zhichen Dong, Yuyu Fan, Xiangtian Li, Chao Yang Distillation can transfer hidden traits from a teacher: a teacher biased by a system prompt can generate semantically clean data, such as numeric sequences, that still makes a student inherit the bias, a phenomenon called subliminal learning, but how the signal accumulates during training has been unclear. The authors propose trait-direction drift as the mechanism: biased generation creates measurable preference gaps in the teacher data, and student-recognizable gaps induce trait-aligned parameter updates during supervised fine-tuning that build up into behavioral transfer. Guided by this, they introduce probe-space corridor regularization, which constrains drift along a calibrated trait direction during distillation, and show it lowers malicious-response transfer from 29.55% to 6.45% with little main-task accuracy cost and consistently suppresses animal-preference transfer in the Qwen setting.

CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs

Maryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi, Martin Takac, Salem Lahlou, Nils Lukas Large language models can reproduce memorized text verbatim, but copyright defenses are typically evaluated under incompatible protocols. CopyShield is a controlled benchmark comparing three defenses at different intervention levels, contrastive decoding at the output level, Direct Preference Optimization (DPO) at the behavioral level, and activation intervention at the representation level, on LLaMA-3.1-8B and Mistral-7B-v0.3 with controlled memorization of five public-domain books, measuring literal leakage, calibrated non-literal leakage, utility, and degeneracy. On LLaMA-3.1-8B, DPO nearly eliminates literal leakage (0.263 to 0.002) but induces paraphrase-loop degeneracy in 58% of QA outputs, contrastive decoding stays nearly degeneracy-free but hits a literal-suppression floor, and activation intervention achieves the lowest non-literal flagging rate by blocking 84% of non-literal queries before generation. On Mistral-7B-v0.3 the output- and representation-level patterns persist while DPO degeneracy falls to 10 to 14%, and targeted non-literal suppression remains an open problem.

Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate

Kaiyan Wen, Shijie Zhang, Lu Yu, Guangdong Bai Text-to-image (T2I) models remain vulnerable to jailbreaks that elicit Not-Safe-For-Work (NSFW) content despite heterogeneous safety stacks combining text filters, image classifiers, and cross-modal detectors, and existing attack studies either target individual filters or query the whole pipeline with aggregate feedback, making it hard to identify the active constraint. The authors introduce the Detection Surface, a geometric framework describing the joint decision boundaries these filters induce, which shows that successful evasion lies in a sparse, non-convex region shaped by cross-layer conflicts where bypassing one filter can increase exposure to another. Building on this, CRACK is a multi-agent debate framework in which an Attack Agent, a Defense Agent, and a Judge Agent iteratively generate prompt mutations, obtain layer-specific diagnostic feedback, and refine strategies with reward guidance. Across multiple T2I models, datasets, and safety configurations, CRACK reaches attack success rates up to 99.63% under composite defenses with fewer queries than prior methods while preserving semantic fidelity.

One Prompt Is Enough: Watermark Laundering Through Foundation Image Models

Jidong Yang, Qi Li, Wei Zong, Yang-Wai Chow, Willy Susilo, Huaike Yu et al. cross-listed Invisible image watermarks are typically stress-tested against fixed perturbations like compression, blur, noise, and cropping, but public foundation image models open a different attack: submit the watermarked image with a single reconstruction prompt and receive a visually faithful output whose watermark no longer decodes reliably. The authors formalize this as watermark laundering and measure it with a joint payload-fidelity profile combining bit error rate with visual and semantic preservation across six OpenAI and Google image editing models, three watermarking schemes, and 1,800 reconstructed outputs. OpenAI models produce the strongest payload disruption across all evaluated schemes, while Nano Banana 2 shows DwtDct remains vulnerable even under high-fidelity reconstruction. Prompt ablations show no removal-oriented instruction is necessary, so the effect comes from the reconstruction pathway itself, motivating foundation-model reconstruction as a missing robustness condition in watermark evaluation.

Position: Privacy Is a Claim, Not a Property of Synthetic Data

Jiachen Zhao, Antonia Januszewicz, Taeho Jung Synthetic data is widely used in privacy-sensitive machine learning, but its privacy protection is increasingly assumed from the generation process itself rather than argued as residual inference risk under stated assumptions. Through an empirical analysis of recent publications across major ML venues, the authors show that synthetic data is frequently deployed in privacy-sensitive settings without explicit threat models, inference risks, or falsifiable privacy claims, leaving assurance implicit, hard to verify, and unevenly distributed, with rare and minority records most exposed. They argue that privacy should be treated as an explicit, evidence-based scientific claim and recommend venue norms requiring privacy assertions to be clearly scoped, testable, and contestable.

The Constitutional Coverage Trilemma in AI Governance

Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly, Moustapha Cisse Each deployed frontier model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity, and the authors ask whether the range of such constitutions on offer covers what people actually want. They pair a paraphrase-controlled audit of the as-shipped default behaviour of 23 frontier LLM archetypes with a pairwise-tradeoff survey of 1,649 US participants on the same instrument. Demand spans all five values with the largest constituency under one third, whereas the 23 shipped archetypes cover only about 2% of the demand space and none put helpfulness or autonomy first, leaving 37% of users without a matching model, and version-over-version drift within families moves further away from autonomy. A sparse menu of two archetypes prioritizing honesty and autonomy would cut mean regret by 47% relative to the whole current frontier, which the authors formalize as a budgeted-pluralism trilemma.

VerTox: Verifiable Reward-Guided Corpus Poisoning Against Neural Ranking Models

Zhiqi Huang, Vivek Datla, Zhichao Xu, Puxuan Yu, Vivek Srikumar, Alfy Samuel Neural ranking models sit at the core of search and retrieval-augmented generation (RAG) pipelines, and their robustness to fluent machine-written spam is poorly characterized. VerTox casts corpus poisoning as reinforcement learning with verifiable rewards (RLVR), shaping the reward to jointly maximize ranking distortion and factual corruption while finetuning compact language models into adversarial document generators. The attack reaches near-perfect success rates across major neural ranking architectures and a proprietary commercial embedding model, producing low-perplexity documents that are hard to detect and that measurably degrade a downstream RAG application.

IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals

Md. Atabuzzaman, Christian Alexander, Chris Thomas cross-listed Factual correctness in large vision-language models (LVLMs) is usually policed by external verifiers or generation-time confidence scores, which add dependencies and still miss outputs that are confidently wrong. IntroConformal is a training-free Conformal Risk Control (CRC) framework whose conformity scores come from the model's own internals: layer-wise semantic stability of hidden-state representations, and a stronger verification probability capturing the model's self-administered judgment on whether a claim is factual. Across several LVLM architectures it holds the finite-sample, distribution-free risk guarantee while abstaining substantially less often than verifier-based baselines, with comparable or better claim-level discrimination.

When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

Yitong Guo, Xiaoyi Chen, Siyuan Zhang, Xiaofeng Wang, Haixu Tang cross-listed Fine-tuning a large language model on entirely benign data reliably erodes its refusal behavior, usually blamed on gradient conflict between objectives. An alternative account is offered in terms of Fisher information geometry: the safety-relevant Fisher matrix is low-rank, alignment flattens the safety landscape while leaving an output-routing pathway intact, and after only 100 benign fine-tuning examples that pathway is selectively re-sharpened in output-side MLP modules. The routing view explains the asymmetry — attack success rates can collapse safety entirely while general utility degrades only mildly — and why a handful of safety examples restores refusals, since the internal safety representations were never destroyed. LoRA and ASAM delay early collapse by suppressing output-side sharpness, but their protection weakens at larger fine-tuning scales.

Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents

Xiaofang Yang, Ziqi Miao, Dianbo Sui, Jing Shao, Lijun Li cross-listed Agents that load reusable skills as persistent runtime context give a malicious skill a durable channel for steering later actions — leaking secrets, corrupting code, or staging exfiltration only once a concrete task makes the action look useful — so pre-install vetting is not enough. Defense-as-Skill makes the runtime guard itself an installable, inspectable, editable skill: SkillSonar runs alongside untrusted skills, checks sensitive actions against the user's task boundary, and routes each to allow, replan, or confirm without modifying the agent runtime. Evaluated on SCOPE-R, a new task-conditioned dataset of 206 attack-confirmed malicious instances across 6 risk families plus 43 benign tasks, and improved by a Monte-Carlo tree search that evolves the on-disk guard from rollout feedback, it cuts in-distribution attack success from 0.482 to 0.104 and out-of-distribution from 0.606 to 0.115 on GLM-5. Protection transfers across victim models, held-out risk families, and external benchmarks, and holds up against adaptive attackers on Claude Code and OpenClaw.

Mechanism Design for Alignment and Control

Dirk Bergemann, Andrew Koh, Stephen Morris cross-listed The framework treats mechanism design for AI agents whose alignment, meaning their preferences, and capabilities, meaning feasible actions and information, are both unknown, so mechanisms that make such agents act on our behalf must incentivize honesty and obedience together. A one-sided imitation structure, in which capabilities can be concealed but not counterfeited, yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. The framework is applied to stylized examples of sandbagging, where a capable agent pretends to be less capable, an alignment-interpretability trade-off in which the two are substitutes in the instrument but complements in value, discipline via peer scoring, coupled rewards that induce competition among agents, and scalable oversight with reward shaping.
4 more specialized papers

Other 42

Modelpedia: A Catalog of Model Findings for the Meta-Science of AI

Franciszek Bernat (Centre for Credible AI, Warsaw University of Technology), Dawid P{\l}udowski (Centre for Credible AI, Warsaw University of Technology), Micha{\l} Jan W{\l}odarczyk (Centre for Credible AI, Warsaw University of Technology) et al. Findings about how foundation models behave and fail are published faster than the community can organize them and end up scattered across papers, blogs, and technical reports. Modelpedia is an automated, LLM-assisted framework that extracts findings about models from published papers, links each to the model, dataset, method, and concept it concerns, and aggregates them into a searchable public catalog. Applied to accepted ICLR 2024 and 2025 papers, the prototype extracts over a thousand findings, and the authors treat the catalog itself as an object of study to run a meta-analysis of how the community investigates models, inviting contributions to the open resource.

Superposed Latent Autoencoder

Quanling Zhao, Jiaying Yang, Tianqi Zhang, Ziyang Hao, Fatemeh Asgarinejad, Flavio Ponzina et al. Autoencoders usually meet tight latent-memory budgets by shrinking each latent, which sacrifices representational capacity. The Superposed Latent Autoencoder (SLAE) instead keeps latents wide and shares storage: it transforms latents into storage-friendly codes, binds them with randomized keys, superposes several codes into a single memory tensor, and learns to recover each latent before decoding, replacing an irreversible dimensional bottleneck with structured interference that can be suppressed. Across CIFAR-10, CIFAR-100, SVHN, STL-10, and Tiny ImageNet, SLAE reduces reconstruction error by up to 56% over conventional autoencoders at matched storage, and the preserved information lifts downstream classification by up to 16.79 percentage points under the same memory budget.

Births are difficult to predict even with rich survey and full-population register data

Elizaveta Sivak, Emily M. Cantrell, Thomas Emery, Javier Garcia-Bernardo, Flavio Hafner, Kasia Karpinska et al. Major life events are notoriously hard to predict, and it is unclear whether that reflects weak theory, data, and algorithms or the large role of chance. A data challenge had 147 researchers predict whether Dutch residents aged 18 to 45 would have a child within three years, using either rich survey data or full-population registers, with methods ranging from logistic regression to transformers and a large language model. Predictions were only moderately accurate (best F1 of 0.59 on register data and 0.76 on survey data), advanced models did not beat classical ones, and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy yields a predictive ceiling of roughly 0.86 to 0.96 F1, so observed performance falls short of what data and methods could achieve, while chance in reproduction alone still imposes a non-trivial limit on predicting individual lives.

Recent Developments in Transformer Inference Deployment on FPGA Platforms: A Survey

Arjan Blankestijn, Uraz Odyurt, Amirreza Yousefzadeh Deploying Transformer models for inference demands throughput, latency, and energy efficiency that Central Processing Units (CPUs) and Graphics Processing Units (GPUs) do not always deliver, and Field Programmable Gate Array (FPGA) platforms offer implementation flexibility, energy efficiency, and suitability for on-site deployment as an alternative. The authors conduct a systematic literature review of recent Transformer inference work on FPGAs, extracting the preferred implementation and optimisation techniques and organizing design choices, trends, and optimisation methods into a taxonomy. The review is intended as a guide for both academic and industry practitioners choosing how to deploy Transformers on reconfigurable hardware.

Optimizing Byzantine Node Placement in Decentralized Federated Learning

Edoardo Gabrielli, Gabriele Tolomei Security studies of decentralized federated learning focus on how Byzantine participants behave and largely ignore which participants get compromised, even though aggregation runs over a communication graph where placement decides how far malicious influence spreads. Treating placement as an explicit adversarial choice under a fixed compromise budget, Byzantine Placement Influence (BPI) is a set-level measure derived from the gossip dynamics that quantifies honest nodes' cumulative exposure to Byzantine sources over the training horizon, accounting for weighted multi-hop propagation and interaction among compromised nodes instead of relying on centrality heuristics. Across six heterogeneous graph families, untargeted model poisoning, and backdoor attacks, BPI-guided placements consistently identify the most damaging configurations, and remain effective when Byzantine-robust aggregation breaks the linear gossip assumption.
37 more specialized papers

Theory 42

Dense Weak Hiding: Closing Complexity Gaps in Nonconvex and PL Finite-Sum Optimization under Individual Smoothness

Yuxing Peng, Zhiqing Tang, Weijia Jia cross-listed For nonconvex finite-sum optimization under individual smoothness, known algorithms use O(n + √n·ΔL/ε²) incremental first-order oracle calls while the best lower bounds fell short by a factor of √n, leaving the optimal complexity open. A matching randomized lower bound closes that gap — fixing the minimax oracle complexity up to universal constants under both individual and mean-squared smoothness — and companion bounds settle the global Polyak-Łojasiewicz regime, showing restarted PAGE is optimal where the standard guarantee was not tight. The technique, dense weak hiding, spreads each hidden direction across components via a fixed sign table so an individual queried row leaks little while the exact row average preserves the signal, with a bounded radial map and a smooth gate keeping unopened links invisible to values and gradients alike.

Recursive Criticality of AI Self-Improvement

Mikhail Burtsev A dynamical model asks when AI systems used in AI research and development produce self-amplifying capability growth, balancing baseline research productivity, the strength of recursive feedback, and the rising difficulty of further progress. The central quantity is a recursive reproduction number whose value above one marks a regime where improvements compound across development cycles, and below one one where they are damped; the threshold does not correspond to any particular capability level, so a system can enter the amplifying regime before acceleration is visible, while rapid progress can also occur without amplification. Extending the model to multiple actors shows improvements shared between organizations can make the overall research ecosystem self-amplifying even when no individual organization is, and the framework names measurable properties — feedback strength, propagation into successor systems, cycle duration, difficulty growth — that distinguish amplification from fast progress with other causes.

Rock, Paper, Scissors, ... Dynamite - A Model of Disruption from New Technologies

Andrew J. Lohn cross-listed To study how a disruptive new capability reshapes a competition, the authors add a versatile "Dynamite" move to Rock-Paper-Scissors and solve for equilibrium play. Giving Dynamite to only one player raises that player's win probability from 50% to just 55.5%, and it gets played rarely; the advantage shrinks further when the game is expanded beyond the original three moves. The analysis also surfaces mechanisms by which existing moves become strategically unplayable or obsolete, which the authors offer as an intuition pump for how raw capability translates — or fails to translate — into value.

Exact Global MCMC with Denoising Diffusion

Mitch Hill cross-listed Sampling from complex high-dimensional unnormalized densities is hard for local Markov chain Monte Carlo (MCMC) methods that get trapped in modes. The observation here is that running a forward then reverse diffusion process defines a Markov chain preserving the target distribution when the denoiser is ideal, and that this can be made exact for any denoiser by adding a Metropolis-Hastings correction whose acceptance ratio uses the densities of the forward and reverse paths of a discrete-time stochastic differential equation approximation. Denoising Diffusion Monte Carlo trains a standard denoising diffusion model on locally convergent Metropolis-adjusted Langevin samples and composes the resulting global path proposal with a local sampler, achieving high acceptance rates for global moves across a range of complex targets and offering preliminary evidence that diffusion training's scaling behavior carries over to exact sampling.

How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks

Arnol Manuel Fokam, Fasseu Sieyondji Akpevwoghene, Edem Fiifi Dawson Prior analyses of how much memory a linear recurrent neural network builds during training assumed uncorrelated inputs, whereas real sequences are correlated. Solving the learning dynamics exactly for correlated inputs shows the entire effect of correlation collapses onto a single cost attached to retaining the past, recovering the earlier result when inputs are uncorrelated and growing under positive correlation. Three consequences follow: correlation reshapes the whole course of learning, with memory building, overshooting, and being partly removed before the network settles on retaining less; memory switches off at a threshold set by one number — how much each input resembles the one immediately before it — independent of sequence length or longer-range correlation; and given a single spare hidden dimension, training spontaneously builds a feedthrough path that passes the current input straight to the output and remembers nothing. The analysis turns one property of the input into a prediction of whether memory is learned at all, and explains why correlated data turns recurrent networks into change detectors.

Independent Reinforcement Learning in Discounted Markov Games

Asrin Efe Yorulmaz, Ugur Aydin, Tamer Basar cross-listed The setting is radically uncoupled learning in discounted general-sum Markov games, where each player updates from its own rewards with no knowledge of the others. Assuming the exponential time hypothesis for the complexity class PPAD, the authors prove that for every fixed discount factor no polynomial-time independent-learning algorithm can compute inverse-polynomially accurate coarse correlated equilibria. They then give what appears to be the first radically uncoupled algorithm with sub-exponential convergence to coarse correlated equilibria in this class of games without structural assumptions: a layered variant of optimistic mirror descent with an increasing step-size schedule, in both full-feedback and partial-feedback versions.

Disciplined Bilevel Programming

Hao Zhu, Joschka Boedecker cross-listed Bilevel optimization models hierarchical decisions where one problem is nested inside another, but using existing solvers demands manual reformulation by an expert. Disciplined bilevel programming (DBLP) lets users write optimistic bilevel problems close to their mathematical form, then automatically canonicalizes a convex lower problem into conic form and builds an equivalent single-level reformulation via conic Karush-Kuhn-Tucker conditions, relaxing the complementarity constraint and solving a sequence of smooth nonlinear problems through gap continuation. The implementation, BLVPY, extends CVXPY and lets users state and solve bilevel problems in a few lines of code without bilevel modeling expertise, demonstrated across several application domains.

Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets

Cheung Hao Lee, Patrick Wong A service fronting several language models must choose one per request while respecting workload-level budgets for compute, latency, memory, or cost, and both the request mix and the model lineup drift after launches, fine-tunes, and quantization changes. The setup is formulated as nonstationary sparse contextual routing under multiple knapsack constraints, with an optional shadow-audit stream that scores a small fraction of prompts on several models. DRS (Drift-Aware Sparse Routing) estimates reward and resource use from a rolling audit window, routes on pessimistic reward and optimistic cost, updates resource shadow prices online, and applies a hard budget meter before committing. Regret against a paced dynamic benchmark is bounded in terms of sparsity, audit rate, window length, and drift, recovering the standard O(√(sT/ρ)) rate when there is no drift and a T^(2/3) adaptation term when there is.

Verdict Instability of OOD Scores under Reference Resampling

Donghoon Lee, Shinjin Kang Post-hoc out-of-distribution detectors are fitted on a finite reference set, so every score is an estimate and a different reference sample would move some verdicts; the authors measure that movement as the bootstrap standard deviation of the score, called verdict instability, and derive a closed form with no fitted parameters. Instability turns out to be the within-class dispersion of the assigned class along the query's direction divided by the square root of that class's reference count, and the count term is identifiable only under class imbalance. Because far-OOD queries lie along the low-variance directions of anisotropic embeddings, every distance-based score tested assigns its most extreme values to its most reproducible verdicts, and only estimators of local dispersion carry the sign practitioners expect. A single label-free correlation predicts that sign for any score, and abstention driven by a wrong-signed score is worse than abstaining at random on every dataset tested.

Denoising Diffusion Generative Models Secretly Calculate Attentions

Farzan Haddadi, Leila Monfared, Ebrahim Rezaii, Mohammadreza Malek-Mohammadi, Pejman Zakalvand, Narges Mokhtari Denoising diffusion models dominate image generation while attention-based transformers dominate language modeling, and the authors argue the two families share more than they appear to. They show that diffusion models inherently compute an attention-like operation close to the transformer's, and draw a parallel between auto-encoders and attention-based models, suggesting the designs can be swapped depending on practical requirements. Using this equivalence, they reformulate the diffusion framework into a simplified attention-based image generation algorithm and report that it achieves comparable performance with significantly less training effort and compute.

When Metropolis and Hastings Meet Bradley and Terry: Exact MCMC From Preference Voting

Ariel Smogorghevski, Nir Rosenfeld, Yaniv Romano Sampling from a generative model's distribution conditioned on a semantic property is a natural fit for Metropolis-Hastings (MH), but MH needs exact pointwise density ratios, which are unavailable in generative settings where only pairwise comparisons from humans or model judges are cheap. Pref-MH observes that the MH unnormalized density ratio equals the preference odds under the Bradley-Terry (BT) choice model, and develops an accept/reject rule that works from sampled binary judge feedback while still provably converging to the target distribution. The authors show that for a fixed proposal kernel and comparison budget, the rule is Peskun-Tierney optimal among exact reversible acceptance rules of this class. Experiments on text generation and molecular design with LLM judges, and on image generation with vision-language model judges, demonstrate practical conditional sampling from comparative feedback alone.

The Multiple Timescales of Gradient Descent on the Edge of Stability: A Perturbative Derivation of the Central Flow

Rapha\"el Berthier The central flow of Cohen et al. is an empirically accurate continuous-time model of gradient descent at the edge of stability, but its original derivation was heuristic. The authors propose a perturbative regime in which the loss decomposes as a main term plus a small perturbation, treat gradient descent as a singularly perturbed dynamical system, and apply the classical method of multiple scales to expand the dynamics. Three timescales emerge, fast oscillations along the sharpest direction, an intermediate self-stabilization mechanism, and slow motion along the minimizers of the main term, and the central flow appears as the leading-order term of the expansion with self-stabilization at the next order; the analysis also computes the slow drift of fluctuation energy and explains why fluctuations persist when several eigenvalues sit at the edge of stability.

Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras

Jiming Feng, Junliang Li A sparse subset of Transformer attention heads has effective output-value (OV) operators that nearly close under composition, with the operator squared approximately equal to a scalar multiple of itself, a property the authors call scaled idempotence. Across six pretrained models from 2.8B to 235B parameters, 3.98 to 8.00% of heads reach a squared-closure alignment of at least 0.9, while no matched within-layer O/V mismatch does, and an exact principal-coordinate factorization separates within-support transport from read-write return geometry. Scrambling only the orientation of the core factor while preserving singular values, norms, factor spans, and principal angles collapses median closure from 0.336 to about 1e-4 across 7,304 heads in nine multi-head and grouped-query attention models, showing the property is a trained orientation within broadly available geometric capacity. Under exact value sharing, headwise closure extends to a right-action algebra across heads, verified approximately in seven models.

On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study

Chathurika S Abeykoon, Mathias Nthiani Muia, Mallory Goldstein cross-listed Generative augmentation is a standard remedy for class imbalance, but its effect on downstream generalization lacks a theoretical account. Framing augmentation as a distribution-mixing process shows the resulting risk distortion is controlled by augmentation strength and the class-conditional Wasserstein discrepancy between real and generated data, and a Rademacher-complexity bound makes explicit the trade-off between hypothesis capacity, augmentation intensity, and generative fidelity. Experiments with Conditional GAN and Conditional WGAN-GP on binary and multiclass imbalanced tasks confirm that CWGAN-GP achieves lower Wasserstein discrepancy, but higher generative fidelity does not reliably translate into better classification, with classical oversampling often remaining competitive.
28 more specialized papers

Multimodal 40

SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces

Ranjit Raut, Aarav Subedi, Sagun Rai, Sudan Jha Architecture drawings, flowcharts, and pipeline schematics often carry more information than the surrounding prose, yet no public corpus pairs that figure type with the captions, context, questions, answers, and reasoning steps needed to train a vision-language model on them. SCAFFOLD builds such tuples from arXiv computer science papers using layout detection and PDF parsing plus an AI-assisted question-generation stage, yielding 157,387 pairs across 29,887 figures from 3,058 papers in the largest split, with 37K and 12K subsets alongside it. Baseline experiments use the 12K subset on Qwen2.5-VL-3B-Instruct.

Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy

Yi-Cheng Lai, Hen-Hsen Huang Multimodal LLMs will let conflicting text override what an image plainly shows, a failure the authors term multimodal contextual sycophancy and probe with a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text. The key manipulation is where the information boundary sits around a context-blind visual witness: on abnormal images paired with false Gemini-generated text, GPT-5.1 scores 7.9% when conditioned jointly, 49.7% when the context-blind witness report is scored directly, 63.7% in a two-call witness-arbiter pipeline that shows the witness the text, and 84.2% under System-2 Visual Arbitration, which withholds the text from the witness. Across six models that last setup beats the direct witness report by 19.7 to 44.1 points with all paired confidence intervals excluding zero, though the best boundary is model- and source-dependent — text helps some models, and a GPT-4o-regenerated subset reorders the conditions.

Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy

Shmuel Berman, Jia Deng Memory benchmarks for language and vision-language models (VLMs) usually report accuracy over long text or video, which reveals nothing about how much computation an answer costs or whether the system knows when to abstain. ECCBench adds three axes to capacity: efficiency measured in FLOPs needed to answer from memory, compression (whether compressible inputs are remembered more accurately or more cheaply), and calibration (abstaining in proportion to uncertainty and the cost of an error). Pretrained VLMs turn out to compress their memory over text but not over video, and are poorly calibrated on both, while several non-Transformer memory backbones achieve better compression-calibration tradeoffs than RoPE Transformers, suggesting them as components for long-horizon agents.

Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation

Ruotong Wang, Zihao Zhu, Siwei Lyu, Xin Tao, Baoyuan Wu cross-listed Multimodal large language models used for video moderation miss a compositional failure: a video whose individual parts are all benign can still convey harmful meaning as a whole, a phenomenon the authors call Distributed Implicit Harm and study along two axes — harm distributed across temporal visual segments and harm arising between audio and visual streams. Because such videos are absent from safety datasets and resist retrieval by keywords or local visual cues, the authors built a multi-agent synthesis pipeline that composes benign components into harmful scenarios, producing over 9,000 videos with explicit reasoning annotations. Benchmarking more than 30 models, including frontier proprietary systems, revealed consistent deficits on both axes: models often judge each component correctly in isolation yet fail to see the meaning emerging from their combination, and the same failure appeared on real social-media videos collected manually.

CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction

Zhengxu Tang, Guofeng Cui, Ziyu Gong, Xiaozhou Zhang, Ruifeng Deng, Chengzhi Qi et al. cross-listed Long-tail driving failures are usually treated as rare-object recognition problems, but the decision-relevant question is how an unusual object constrains the ego vehicle's feasible actions. The authors define decision-level driving affordance prediction, mapping a front-view image, ego-motion history, and navigation command to a structured longitudinal-lateral meta-action, and release CoLT-Drive, a 3,536-sample counterfactual benchmark that inserts rare objects into otherwise fixed scenes. Their adaptation method KPA combines perception-to-decision prompting, spherical-interpolation expert merging, and RegMoE, a regime-aware mixture of low-rank adapters, to add task capacity without erasing pretrained open-world knowledge. On CoLT-Drive it reaches 60.8% action-pair accuracy versus 50.3% for the Qwen3-VL-2B baseline and 32.4% for plain LoRA supervised fine-tuning, which degrades sharply.

Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts

Athulith Paraselli, Etha Tianze Hua, Ellie Pavlick Vision-language models must sometimes choose between information supplied in context and facts memorized during training, and the resolution turns out to depend on modality. Models tend to follow in-context information for entities named in text but fall back on memorized facts for entities shown in images. The authors attribute this to late representational alignment across modalities: resolving a visual entity takes longer, so the model's usual factual-recall mechanism is never suppressed in time, yielding parametric answers. Chain-of-thought prompting does not close the asymmetry, though adding more visual information to the context does shift behavior.

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo et al. Speculative decoding speeds up generation without changing outputs, but on vision-language models it has been stuck in a loop: the drafter is autoregressive so it must stay small, a small drafter cannot process the image every step, and a vision-starved drafter fails exactly where the image would make text predictable. GLANCE breaks this with a block-diffusion drafting head that reads the target model's already-fused vision-language state, so images cost the drafter nothing, and fills an entire block in a single forward pass, with a wide candidate tree verified in one target pass and every audited prompt reproducing greedy decoding exactly. Under one engine and round budget it decodes up to 2.93x faster than autoregression using one draft pass per round where the production EAGLE3-VL head takes eight, and accepts blocks 2.7x longer than an EAGLE-3 head trained on the same data. The authors also fit a relationship between accepted length and the target's next-token entropy whose slope steepens with grounding, holding across five tasks and multiple targets while predicting where free-running text still favors chain drafting.

(V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement

Zach Studdiford, Kanishka Misra Whether language models learn abstract grammatical rules or merely track lexical co-occurrence is hard to settle with text alone, since distributional cues like is/are and this/these give grammatical number away. The test here uses vision-language models (VLMs): new nouns are taught by adding fresh embeddings and updating only those, comparing a condition where singular versus plural is signalled purely by the image against one where text disambiguates it. Across behavior, representational dynamics, and causal interventions, the models show non-trivial cross-modal generalization in both exposure conditions, and the internal mechanisms treat linguistic and extra-linguistic cues in similar ways, which the authors read as abstraction-compatible behavior rather than surface pattern matching.

Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict

Jungyeon Lee, Yejin Yoon, Taeuk Kim The same evidence can be handed to a multimodal model as text, as a rendered image of that text, or as both, and it is unclear whether these surface forms are treated equivalently when the evidence contradicts what the model already believes. Testing 13 multimodal large language models on two datasets under knowledge conflict, the authors find that models more readily accept contradicting context in image form than in text form, and that when conflicting text and image appear together the winning modality is essentially arbitrary, shifting with input order, model, and dataset. The instability degrades multimodal retrieval-augmented generation and is exploitable by adversarial attacks; of the mitigations tried — prompting, activation steering, supervised fine-tuning, and direct preference optimization — only supervised fine-tuning helps moderately.

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

Yijun Chen, Yaqi Zheng, Yanya Li, Boyi Xiao, Buqiang Xu, Shuofei Qiao et al. Memory systems for long-video question answering usually store captions, frames, transcripts, and graph facts as separate fragments, forcing the language model to reassemble cross-modal and temporal alignments at inference time when context is scarce. EM^2Mem instead binds heterogeneous evidence to event anchors while memory is being built, so each event-indexed cell already carries aligned multimodal records, temporal context, graph relations, semantic facts, and provenance. Across three long-video QA benchmarks it gains 2.0, 2.4, and 3.7 accuracy points over the strongest memory baseline and 7.0 points of strict event-level top-5 evidence recall, while cutting per-query latency 4.67 times and total inference tokens by 63.66%.

VoiceLongMemEval: Do Assistants Remember How You Sounded?

Ramit Pahwa, Parivesh Priye, Apoorva Beedu Long-horizon memory benchmarks for assistants test what was said across sessions but not how it was said, ignoring emotion, prosody, and voice events. VoiceLongMemEval makes every answer depend on paralinguistic metadata attached to conversational turns, with a three-stage adversarial gate ensuring a strong text-only language model fails on each item. Evaluating frontier and open-weight models exposes a pervasive affect gap: supplying the paralinguistic metadata as text raises accuracy by 0.09 to 0.38 (reaching 0.61 to 0.69 with evidence hints), audio-native models recover some of the signal directly from speech (0.354 to 0.412 versus 0.325 blind), and standard speech-recognition pipelines discard it entirely.

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

Haoyuan Shi (Hunyuan, Tencent), Mingtao Chen (Hunyuan, Tencent), Shuo Jiang (Hunyuan, Tencent) et al. Commercial short-drama production runs through script, storyboard, keyframe images, shot-level video, and final assembly, yet existing benchmarks score only the video-generation stage using hand-authored inputs rather than actual upstream outputs. DramaChain Bench evaluates every stage with a shared set of five axes expanding into 63 leaf dimensions, backed by 5,785 items each scored by three professional annotators, producing 17,488 scores and 255,925 attribution records with defects localized in space and time. The human annotations confirm that upstream defects cascade downstream, so final episode quality is not determined by video generation alone, and an agentic automatic judge that gathers evidence over multiple rounds reproduces the human model ranking at a mean PLCC of 0.918.

SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task

Qiming Bao, Ne\c{s}et \"Ozkan Tan, Siyuan Wang, Mark Gahegan The NTCIR-19 SciClaimEval task asks systems to verify scientific claims against the tables and figures of a paper, and the entry described here benchmarks eleven frontier and open multimodal models under one per-sample protocol instead of tuning a single system, then adds light post-processing. The team placed first in three of four evidence-category and subtask combinations, with GPT-5.5 and Claude Fable 5 leading both subtasks and Claude Opus 4.8 and Gemma-4-31B beating the strongest public baseline. The largest gain came from the data's structure rather than model choice: a leak-free pair prior raised Subtask-1 pair accuracy from 72.2 to 93.5, more than any model swap or ensemble. A case-by-case audit attributes most remaining errors to label noise or label-mapping swaps, and documents a measurement leak in which released file ordering encodes the answer.

Controllable Image Captioning with Prompt-Conditioned Scene Rewards

Jongyeop Hyun, Taeyoung Kim, Hyounghun Kim cross-listed Large vision-language models write fluent image descriptions but give users little control over whether a caption emphasizes attributes, relations, or particular regions. FoCUS adds a prompt-conditioned reward: generated captions are parsed and aligned to scene-graph components such as objects, attributes, and relations, those components are weighted — negatively when the user asks to avoid something — according to the control prompt, and the model is optimized against the resulting objective with GRPO, backed by a stricter object validity threshold and reasoning-based verification for attribute and relation scoring. The authors also introduce SCoPE, a benchmark of contrastive Include/Avoid constraints that scores both coverage of requested content and suppression of out-of-scope content, and report improved controllability and fine-grained caption quality on two VLM backbones without degrading general captioning.

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj cross-listed Audio language models are meant to understand speech, but it is unclear whether they represent how something is said rather than only what is said. Using the Expresso dataset of controlled speaking styles, the authors trace paralinguistic information through Whisper-large-v2, Qwen2-Audio-7B-Instruct, Qwen2.5-Omni-7B, and Chroma-4B with centered kernel alignment, leave-one-speaker-out linear probing, open-ended tone prediction, and a content-prosody leakage metric. All four models encode speaking style strongly in the top third of the audio encoder, yet that information is consistently degraded before it reaches the output: the projector reshapes geometry without destroying style, while decoders split into content-driven behaviour, where predictions track the text, and acoustic-driven behaviour, where they vary with delivery.

Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis

Minsik Choi, Geewook Kim, Young Geun Kim Turning a pretrained language model into a vision-language model erodes its text ability, with damage concentrated on tasks graded by strict output formats such as instruction following and strictly parsed chain-of-thought answers. The authors attribute this to attention-sink corruption: visual fine-tuning perturbs the early token position that absorbs a large share of attention probability, and how well a backbone preserves that sink predicts how much capability survives. They introduce Sink Strength, a single scalar computed on the base model in seconds on one GPU that tracks relative post-adaptation degradation across six model pairs without any vision-language training, and report two negative results: injecting QK-RMSNorm after pretraining does not reproduce the protection of native QK-RMSNorm, and several off-the-shelf weight-merging recipes fail to recover the lost ability.

Solaris: Towards Interfaces That Are Generated, Not Coded

Yuval Alaluf, Omri Avrahami, Guy Bukchin Leshem, Michal Geyer, Kfir Goldberg, Elad Richardson et al. cross-listed User interfaces are normally specified ahead of time through code, which fixes their appearance and behaviour to a predefined set of states. Solaris is an interface world model that skips that intermediate representation entirely: it treats mouse input as a conditioning signal and autoregressively generates each frame of the interface, with few-step distillation and training on its own outputs used to hold visual coherence at interactive speeds over long sessions. A separate language model interprets user intent and specifies how interactions should change the environment, splitting high-level reasoning from visual rendering so that interactions which were never programmed can still be carried out.

Visual Attention Faithfulness in Vision-Language Models is Heterogeneous

Xurui Song, Weishi Wang, Zhongqi Yue, Kuluhan Binici, Tao Bai, Hongxin Shao et al. cross-listed Whether attention weights explain a model's reasoning has been argued over for years in natural language processing, but the question is largely untested for the visual side of vision-language models (VLMs). Causal perturbation analysis measuring both comprehensiveness and the sufficiency gap of attention-ranked visual tokens finds that faithfulness is not a single property of a model but splits into three processing modes: Faithful-Sufficient, where the top-k attended tokens are both necessary and sufficient; Faithful-Distributed, where they are necessary but wider context is still needed; and Non-Focal, where no localised region is individually necessary even though visual input remains essential. Human-annotated ground-truth regions satisfy comprehensiveness in only about 60% of cases relative to model attention rankings, and the mode a model falls into varies systematically with architecture and task across VQAv2, VRDU, and ChartQA.

Towards Generalizable Visually Grounded Exploration of Household Devices

Linhao Zheng, Zeming Liu, Wangke Chen, Li Zeng, Wanxiang Che, Heyan Huang et al. Operating an unfamiliar household appliance without a manual requires forming hypotheses, acting, and correcting from feedback, a loop that existing embodied benchmarks sidestep by supplying documents and annotated demonstration trajectories. VGEBench targets what the authors call generalizable visually grounded exploration by building a logic-driven state machine that simulates multi-turn interaction, forcing a vision-language model to reach goals through active perception and feedback-driven correction rather than imitation. Evaluation shows current vision-language models struggle to translate semantic world knowledge into correct physical action sequences and to track device state over long horizons.

Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

Yumi Lee, Harim Oh, Hyoryung Kim, Minji Kim, Eunsu Kim, Hyeseong Lee et al. cross-listed Generating a pathology report from a whole-slide image is held back by the scarcity of paired slide-report data and by the difficulty of turning spatially scattered visual patterns into structured clinical text. A clinically curated Pan-Asia dataset of roughly 10,500 slide-report pairs from five institutions underpins the REG 2025 benchmark, run as a MICCAI challenge, whose submissions span pretrained vision-language models, multiple-instance learning, hierarchical expert models, retrieval-augmented generation, and cross-modal transformers. The analysis finds that using a vision-language model was not itself what separated top entries — structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding were — and identifies recurring failures including numeric hallucination in attribute estimation and diagnostic overspecification that mirrors known pitfalls in routine practice.

The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence

Genpei Zhang cross-listed Reporting aggregate accuracy on a multimodal benchmark quietly assumes the model actually looked at the image, and that assumption is tested by blurring the question-relevant region and measuring how much the next-token distribution moves. Across six vision-language models and three perceptual benchmarks, the distribution barely shifts on 40% to 97% of samples, a phenomenon named the Visual Insensitivity Gap and scored per sample by a Visual Sensitivity Index; the effect belongs to samples rather than models, since the index correlates across architectures sharing nothing but a contrastively pretrained vision tower (grand-mean Spearman rho = +0.40). A linear probe on each model's own vision tower separates perturbed from clean images at 0.72-0.79 accuracy on exactly those samples while the model's top token changes on only 2-11% of them, locating the failure in the encoder-to-language-model handoff rather than in perception, and the index works best as a conditional signal for multiple-choice reasoning on capable models (AUROC 0.85-0.87) rather than as a universal abstention criterion.

From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding

Raul Ortega, Jos\'e Manuel G\'omez-P\'erez cross-listed Vision-language models (VLMs) handle natural-image question answering well but struggle with scientific diagrams, which convey functional or relational meaning rather than literal scenes. The authors propose a framework that extracts domain concepts from science-curriculum terminology, synthesizes atomic facts, retrieves relevant diagrams from the web, and generates captions and multiple-choice questions as multimodal supervision. The resulting SciGram dataset holds over 194K diagrams and 1.4M visual instructions across life, earth, and physical sciences, and despite noisy web data and synthetic labels, models fine-tuned on it match or outperform state-of-the-art VLMs while using fewer training instances on TQA, ScienceQA, and AI2D. Augmenting LLaVA OneVision with SciGram sets new state-of-the-art results on diagram question answering, and both the dataset and the models are released.

SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models

Shiyu Li, Zi-Yuan Hu, Shijia Huang, Yanyang Li, Yiwu Zhong, Liwei Wang cross-listed Multimodal large language models pay a heavy compute cost for long visual token sequences, and existing pruning methods often mistakenly preserve high-norm outlier tokens that are actually redundant in both feature and spatial dimensions. SinkPruner is a training-free coarse-to-fine framework with a visual sanitizer that filters these high-norm redundancies while easing attention sink and dispersion, followed by a text-guided pruner that keeps tokens semantically aligned with the query. Across twelve image-language and four video-language benchmarks, it preserves 96.5% of LLaVA-1.5 performance and 91.8% of Qwen2.5-VL performance under an 89% token reduction, and the sanitizer also improves existing pruning methods when transplanted.

When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP

Shota Sato, Hajime Kiyama, Tosho Hirasawa, Mamoru Komachi Shrinking the modality gap between image and text embeddings in CLIP is widely expected to improve zero-shot accuracy, yet a smaller average gap often fails to deliver consistent gains. The authors analyze this through the decision structure of zero-shot classification, where accuracy depends on class-wise decision margins rather than average alignment, and use linear correction as a tractable case to show that gap correction can reshuffle relative margins so that predictions collapse onto a small subset of classes, a failure mode they call prediction-level hubness. Experiments across multiple datasets show that accuracy drops under gap correction are consistently accompanied by increased prediction concentration, for both linear and learning-based correction methods, so the authors argue gap correction should be judged by its effect on downstream prediction structure rather than by average alignment alone.

On the Design Fundamentals of Pixel Text Representation Learning

Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang, Yu Rong, Hong Cheng, Hou Pong Chan et al. cross-listed Text-rich visual inputs require encoders that read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders suffer from fixed-resolution pretraining, visual shortcut learning, weak grounding, and poor multilingual coverage. Controlled ablations identify four design principles: variable image resolutions and font sizes as proxies for high-resolution documents, natural image-text pairs to prevent text-only collapse, layout-aware rendering to block pixel-level shortcuts, and a two-stage multilingual curriculum for cross-lingual alignment. These are combined into Pixel Linguist II, a native-resolution vision encoder trained with on-the-fly rendering and unified contrastive grounding over 280M examples, which sets new state-of-the-art results on English, cross-lingual, and multilingual Visual STS and ViDoRe and improves downstream multimodal LLM evaluation. The encoder remains robust under 80% visual token compression, pointing to optical context compression as a use case.

A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation

Hodong Lee, Sanghee Park, Dohoon Ryu, Jungwhan Kim, Junyeob Kim, Soyoon Kim et al. Evaluating an omni-modal foundation model across text, image, video, and audio currently means juggling separate toolkits whose inference engines, prompt conventions, and metric implementations are mutually incompatible. OmniEvaluator connects existing inference engines and curated evaluation libraries at a higher level rather than reimplementing benchmarks, exposing four inference backends, four evaluation frameworks, and over a thousand benchmarks through a single interface, recording every run as a fully reproducible configuration artifact, and feeding results into a shared dashboard for cross-model comparison. A federated mode shares GPU inference servers across concurrent evaluations, and a built-in verifier small enough to run on CPU keeps scores stable across engines and prompts where rule-based scoring fluctuates, matching cost-efficient commercial LLM judges without recurring API cost.

MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval

Debanjan Mahata, Atharva Tendle, Daniel Preotiuc-Pietro, Yong Zhuang, Ozan Irsoy cross-listed Visually rich documents hide key content in tables, charts, figures, and layout that plain OCR corrupts or drops, while ColPali-style visual retrievers fix this with patch-level multi-vector indexes and late-interaction scoring that keep image-derived retrieval on the query-time serving path. MIDR (Multimodal Indexing for Document Retrieval) is a training-free framework that shifts multimodal reasoning to index time, using a multimodal LLM at ingestion to convert rendered pages into verified textual fields indexed with BM25F and optionally fused with dense retrieval. On ViDoRe V3, MIDR Hybrid reaches 0.6219 average nDCG across five English domains, a 23.0% relative gain over BM25, and enrichment lifts BM25 on French documents with English queries from 0.1532 to 0.5448. Across all seven domains it leads ColQwen2.5 on four while using roughly 9x less index memory and 2x lower query latency.

Reliability Challenges in Diffusion Vision-Language Models

Md. Atabuzzaman, Chris Thomas cross-listed Diffusion-based large vision-language models (dLVLMs) offer parallel decoding, bidirectional context, and controllable generation as an alternative to autoregressive (AR) models, but their reliability has not been characterized. The authors benchmark six diffusion models against competitive AR baselines on hallucination and bias across four dimensions. They find that dLVLMs reverse the yes-bias of AR models on binary visual questions, match AR hallucination rates but with degraded linguistic quality, collapse to near-zero accuracy on underrepresented racial groups with opposite-polarity gender bias, and lose accuracy on multiple-choice questions whenever the correct option is shorter than its distractors, a length prior that appears at the first denoising step. Tokens committed late in denoising with low confidence correlate with hallucinated content, a mechanistic signal specific to diffusion generation, and the patterns vary across model families.

EdiTikZ: Scientific Figure Editing from Revision Trajectories

Christian Greisinger, Zhixue Zhao, Steffen Eger Publication-ready scientific figures come from iterative refinement, yet figure editing with vision-language models (VLMs) is largely unexplored, and existing training supervision relies on costly proprietary agent systems or synthetically generated edits. DaEdiTikZ instead mines naturally occurring revision and development trajectories, collecting 391K plausible TikZ edit pairs from arXiv, GitHub, and TeX Stack Exchange and having a VLM infer 781K directed edit instructions from rendered figures and code, paired with a 790-instance human-refined benchmark. Two compact Qwen3.5-based EdiTikZ models (4B and 9B) trained on joint reconstruction and editing followed by reinforcement learning with rendered-fidelity and edit-application rewards put the 9B model above all tested baselines automatically and, across 4,320 human ratings, above GPT-5.6-Sol and on par with Gemini-3.1-Pro.

TempCloze: Can Video-LLMs Identify the Missing Middle?

Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu, Jiahao Meng, Han Chen, Ziyu Wang et al. cross-listed Temporal reasoning benchmarks for video language models are often mediated by language, letting models exploit option wording, answer correlations, or language priors instead of watching. TempCloze reframes the task as a cloze test: given a video's beginning and ending clips, pick the true missing middle from four same-source distractors constructed along Semantic (what should happen), Alignment (when it should occur), and Progression (how it unfolds) axes, with shared scenes and objects suppressing appearance shortcuts, across 1,521 filtered mostly long-take and egocentric videos. Testing 10 proprietary and 21 open-source models identifies Alignment as the primary bottleneck: they recognize plausible content and local event progression but cannot place events correctly in time. Additional TempCloze-Mixed and TempCloze-Hard splits analyze error patterns and sensitivity to candidate order, context direction, visible span, frame density, and test-time scaling.

H3-World: Turning Language Understanding into World Control

Danze Chen, Zeqing Wang, Ziyue Lin, Xingyi Yang, Yeying Jin cross-listed H3-World turns the 33B MiniMax-H3 video generator into an interactive world model by exploiting the observation that large video generators already accept coarse natural-language control of character behavior and camera motion. Each action is represented as a structured combination of character and camera instructions aligned to the corresponding temporal video latents, and temporal attention routing confines each instruction to its intended time interval to reduce control leakage across actions, with no dedicated action modules added. Using only 8,000 gameplay samples, 10,000 LoRA optimization steps, and 0.199% trainable parameters, the system achieves effective character and camera control, preserves generation quality, and generalizes to unseen scenarios.

Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics

Maksim Evdokimov, Matvey Ivanov, Dmitrii Tsiupin, Olga Tsymboi, Anatolii Potapov, Aleksandr Ivanov Extracting structured fields from hundreds of millions of documents a year stays expensive in regulated industries, where bespoke OCR cascades cover few workflows, privacy rules bar external models, and open-source vision-language models (VLMs) that meet quality thresholds cost more to serve than human annotation. The deployed system fine-tunes a Mixture-of-Experts VLM with 35B total and 3B active parameters on in-house production data mixed with open-domain documents selected by a difficulty-aware curation pipeline targeting layout diversity, fact extractability, and cross-model consistency, and it serves heterogeneous workflows through prompting on a single H100. It leads all deployable non-reasoning baselines up to an order of magnitude larger, and a quality-adjusted cost analysis calibrated from production telemetry shows expected costs falling by over 80% against the human baseline and by more than 50% against the best competing open-source model, while larger baselines remain economically unviable.
8 more specialized papers

Vision 17

ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration

Yuchen Bao, Chao Wen, Haowei Wang, Ruoxin Chen, Donghao Luo, Jiahui Zhan et al. Reward post-training of diffusion image generators piles probability mass onto a few reward-favored modes, destroying within-prompt diversity, and existing fixes rely on external signals such as perceptual objectives or text-encoder changes rather than repairing an adapter that has already collapsed. Starting from the observation that online post-training reallocates mass over pretrained capabilities instead of learning new content — so collapse is suppression rather than deletion — ReNFT recalibrates from inside the generator: unconditional probes pick anti-hub prompts where prompt-independent bias is most visible, two mixed rollout routes produce matched counterfactuals from the same prompt and noise, and reward ranking with an adaptive flipping guard assigns pull and push roles for a paired update. On PickScore and GenEval it keeps 98.9% and 99.0% of the reward achieved by NFT while raising DreamSim diversity by 58.8% and 55.0%.

Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You

Salim Khazem, Ibrahim Mohamed Serouis Test-time adaptation normally assumes weights can be updated at inference, which rules out inference-only accelerators, frozen third-party models, and architectures without BatchNorm where standard configurations go inactive. CASTER keeps the network frozen: it stores source class statistics in a discriminative subspace, estimates a class-shared affine transformation from target-batch moments, and analytically transports the source class distributions before classification, with no backward pass, optimizer state, or stored source feature bank. It beats k-nearest-neighbours on identical frozen features in 27 of 28 backbone-dataset settings while retaining a median of 18x less state, but transport is not always safe — on ImageNet-C, where 64-sample batches must cover 1000 classes, unconditional transport loses 21.2 top-1 points. Gating on an empirical residual-to-margin transportability certificate converts an average -3.35-point effect into a +1.69-point gain, though the authors note it does not cleanly separate benign from destructive regimes and is specific to this mechanism rather than transferable to methods like Tent.

Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation

Teresa DiMeola, Charles Walter, Hong Xiao cross-listed Pixel-level segmentation of aerial and satellite imagery for tasks like flood mapping or damage assessment suffers when general foundation models miss classes or small objects in a scene. The proposed pipeline runs a vision-language model (VLM) on a single consumer-grade GPU to supply guidance without retraining: a frozen foundation model labels every pixel, one VLM query selects which classes are relevant to the scene, and a second locates small objects the base model overlooks. Evaluation on four aerial datasets shows consistent gains at each stage where the base model is already competent, plus structured evidence that can be audited independently of the mask.

VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM

Sangmin Song, Sarath Kodagoda, Marc G. Carmichael, Karthick Thiyagarajan, Amal Gunatilake, Kelly Prentice et al. cross-listed Online systems that build open-vocabulary 3D maps typically segment and name an object the moment it is first detected, committing to a decision when visual evidence is weakest. VOIM (Voxel-Grounded Online Instance Manager) instead accumulates soft evidence from off-the-shelf perception models per voxel across views and defers both the instance grouping and the label until that evidence has settled, requiring no training and running from RGB-D or from monocular RGB alone. On ScanNet++ under a matched protocol it reaches 44.07 mIoU versus 32.37 for the strongest online RGB-D baseline, OVO-SLAM, winning all ten scenes, and the same system transfers unchanged to fully monocular input, matching that baseline on Replica. Swapping the region descriptor, detector label prior, and mask source shows the mapping stage rather than the perception models drives the gain, though labelling itself is not real-time because per-class detection over the full vocabulary dominates the cost.

Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking

Orcun Cetintas, Guillem Bras\'o, Tim Meinhardt, Laura Leal-Taix\'e cross-listed Monocular video flattens 3D scenes into 2D projections, and multi-object trackers that rely on image-plane appearance and geometry inherit the resulting depth and spatial ambiguities. PLANET is an end-to-end tracker that first lifts existing 2D tracking datasets into 3D, then forms world-grounded queries by embedding reconstructed scene geometry into the features and positional encodings used to build queries. An auxiliary 3D location prediction task pushes queries to encode object positions during training, and a dual-resolution temporal memory preserves that evidence across longer gaps, yielding state-of-the-art results on three diverse tracking benchmarks.

ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives

Nikos Giakoumoglou, Andreas Floros, Kleanthis-Marios Papadopoulos, Tania Stathaki cross-listed Contrastive self-supervised learning has been overshadowed by generative and self-distillation approaches for pretraining vision transformers. ViTAMINS integrates synthetic hard negatives into unsupervised vision transformer pretraining through simple modifications to existing contrastive frameworks, and is benchmarked on ImageNet, transfer learning, image retrieval, copy detection, and image and video segmentation. The synthetic negatives give rise to representations that explicitly encode semantic content and serve as strong classifiers with gains of up to 11.3% over baselines, and a ViT-B trained this way surpasses V-JEPA with the larger ViT-L while using fewer resources.
11 more specialized papers

Reinforcement Learning 14

Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning

Yu Yuan, Yaoyou Fan, Lili Zhao, Guangting Zheng, Kai Zhang, Lu Pan et al. Reinforcement learning fine-tuning increasingly mixes several reward dimensions — verifiable rules, task evaluators, learned reward models — and collapses them into one scalar with fixed weights. The authors identify a failure mode where that aggregation itself causes reward hacking: static projection maps qualitatively different reward profiles onto the same scalar, pushing optimization toward whichever dimensions are easiest or densest and trapping the policy in suboptimal profiles. AMRP (Adaptive Multi-Reward Projection) reallocates aggregation weights online using relative shortfall, reward volatility, and recent progress, applying pressure to lagging or stagnant dimensions while easing off saturated ones. Across structured reasoning, citation-grounded generation, and open-ended alignment, it improves both reward-profile balance and downstream task performance over fixed and dynamic weighting baselines, and works with GRPO, GDPO, and PPO.

Group Adaptive Clipping Policy Optimization

Sheng Jia, Xiao Wang, Shiva Prasad Kasiviswanathan, Rein Houthooft Group-relative policy optimization for reinforcement learning with verifiable rewards (RLVR) clips the importance-sampling ratio at a fixed boundary for every rollout, which means rare correct rollouts on hard problems get suppressed at roughly the same rate as abundant correct rollouts on easy ones despite carrying much stronger exploration signal. GAPO (Group Adaptive Clipping Policy Optimization) is a drop-in change to GRPO-style methods that scales the clipping boundary with the rollout advantage, motivated by a reverse-KL trust-region argument that higher-signal rollouts deserve more update headroom; it needs no reward shaping and leaves the PPO/GSPO surrogate intact. On Qwen and Llama models it improves both Pass@1 and Pass@k over fixed clipping and advantage-shaping baselines on math and coding benchmarks where base-model pass rates are low.

GeoPAR: Large-Scale Multi-Agent Combinatorial Optimization with Geometry-Guided Parallel Autoregressive Learning

Wenjian Wu, Zesheng Jia, Jiaying Tang, Benyuan Yang, Jin Wang Parallel autoregressive neural solvers let multiple agents pick actions simultaneously in NP-hard routing problems, but they lose quality at scale because local geometry is modeled weakly and conflicting task selections are only repaired after actions are generated. GeoPAR adds a projection-window sparse geometry mechanism that builds lightweight local candidate neighborhoods, sparse edge-biased attention that injects those relations into node representations, and a cache-guided conflict-aware assignment step that suppresses duplicate selections during decoding rather than afterwards. On heterogeneous vehicle routing and open multi-depot pickup-and-delivery problems it improves large-scale zero-shot generalization while substantially reducing the number of rollout steps.

It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning

Runpeng Dai, Kaili Huang, Changsung Kang, Ciya Liao cross-listed Search and advertising retrieval systems increasingly use LLMs to expand queries, but final matching is still handed to a separate retriever. CoGR trains LLMs to generate compact keyword sets for both the query side and the item side, matched through a conventional inverted index so existing infrastructure still works; supervised fine-tuning first aligns the keyword space, then co-evolving reinforcement learning alternately optimizes each side with Group Relative Policy Optimization (GRPO) against the other's frozen index, with the item side rewarded by the counterfactual change its keywords cause in query-side F1. Against 10 sparse, dense, and generative baselines it improves F1 by 10.9% on an internal app marketplace dataset and 36.1% on the public WANDS benchmark.

CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training

Siyuan Li, Xinxin Song, Chen Ruinian, Jingjing Fan, Tingxiong Xiao, Yangen Hu et al. Rubric-based reinforcement learning scores open-ended responses against prompt-specific checklists, but static rubrics get hacked as the policy improves, and existing dynamic rubric schemes suffer from undirected extraction, unreliable hack detection, and unbounded rubric growth. Contrastive Anchor-based Rubric Evolution (CARE) grounds each rubric update in a high-quality anchor response generated by a frontier model, contrasting the highest-scoring rollout against that anchor at every training step. An Adaptive branch reactively repairs reward misspecification while a Chase branch turns quality gaps relative to the anchor into sharper rubrics, keeping the reward discriminative in the high-reward region where over-optimization originates. Trained on WildChecklist-9K with Qwen2.5-7B base and instruct models, CARE reaches state-of-the-art results on Arena-Hard-2.0, InfoBench, and FollowBench, and is the only method whose win rate against GPT-4.1 anchor responses keeps improving across all 300 training steps, with Llama-3.1-8B-Instruct and Qwen3-8B results indicating the gains transfer across model families.

From Base Rollouts to RL Reasoning: A Budgeted Search Perspective

Wenhe Sun, Cunxiang Wang, Zijun Yao, Yixin Cao Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but it is unclear whether it creates capability the base model lacks or merely shifts sampling toward trajectories the base model could already reach. The authors build a Unified Decoding Framework (UDF) that expresses token sampling, beam-like search, tree search, and sequence-level resampling as policies over a shared compute budget, then test on Math500, AIME, GPQA, and IFEval whether an RL checkpoint's pass@k curve can be reproduced by a structured path of base-model operating points, using paired checkpoints from SimpleRL-Zoo. The recovery path follows a Budgeted Operating-Point Transition Rule (BOPTR), a power law relating the base and RL sampling budgets with benchmark-specific exponents, which matches RL curves on Qwen2.5-7B with a transfer error of about 3.4 percentage points and generalizes to ten models across four families and to four benchmarks it was never fitted on. The authors read this as qualified support for an internalized-search interpretation, where much of the measured RL gain under this recipe reflects improved sampling efficiency, and present the rule as a behavioral diagnostic rather than evidence of parameter-level equivalence.

Bandits in Prod: Hyperparameter Optimization at Inference Time

Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine Many production systems can only evaluate a configuration by serving it on live traffic, which is exactly the situation for agent stacks picking a model, retrieval depth, prompting strategy, and decoding temperature without representative validation data. The authors formalize this as online hyperparameter optimization and model it as an infinitely many-armed bandit over mixed, conditional search spaces, yielding IMABO, which pairs any bandit policy over already-sampled configurations with any oracle that proposes new ones. Their restart-free anytime policy IMOSS comes with a proven expected cumulative quantile-regret bound, and paired with a Tree-structured Parzen Estimator, an incumbent-mutation oracle, or a pretrained tabular foundation model, IMABO obtains the lowest cumulative regret across settings ranging from classical machine-learning tuning to configuring LLM-based agents.

Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR

Esther Xin Reinforcement learning with verifiable rewards (RLVR) and benchmark scoring both hinge on an automatic verifier turning free text into a binary reward, and the known figure that one harness accepts only about 94 percent of its own gold answers is an aggregate that says nothing about which answer forms consume the error budget. The authors apply metamorphic testing to the verifiers rather than the models, generating certified meaning-preserving rewrites so every rejection is a provable false negative, then measure per-category rejection over 307,420 verdicts from four widely used verifiers. Self-validation ranges from 53.8 to 95.2 percent on identical inputs, a 41.3-point spread, with two configurations of the same library disagreeing on half of all pairs, and 93.0 percent of in-contract failures for the default LaTeX configuration tracing to whitespace and punctuation such as a trailing period or newline. Separating rejection from execution failure also shows a reference numeric cascade accepting off-by-one wrong answers as a step function of magnitude, from never below 10^4 to always at or above it, because its tolerance is relative.

Provably Safe Sim-to-Real Transfer

Tingting Ni, Maryam Kamgarpour Policies trained in a simulator to avoid the sample cost of real-world reinforcement learning can be suboptimal once deployed, and correcting the mismatch requires real data whose collection is itself safety-constrained in domains like robotics and healthcare. Safe sim-to-real transfer is formulated within the framework of reward-free safe reinforcement learning, with a computationally efficient algorithm that exploits simulator information to provably reduce real-world interaction while keeping exploration safe and yielding a near-optimal feasible policy for any potential reward function. The real-world sample complexity bound quantifies the simulator's benefit as a function of the sim-to-real mismatch.

From Rollouts to Recipes: Self-Contained Post-Training for LLMs

Yifei Li, Lingling Zhang, Muye Huang, Zihan Ma, Jiashuai Liu, Jun Liu Post-training normally applies a single recipe to every sample even though a model's own rollouts reveal that samples sit in different learning states. Self-Routing reads rollout correctness and confidence to send each sample to GRPO, on-policy self-distillation, regularization, or skipping, requiring no external teacher, extra annotation, or additional sampling. On mathematical reasoning with Qwen3 and Qwen3.5 backbones it consistently beats uniform GRPO, uniform on-policy self-distillation, fixed mixtures, and simpler routing baselines, and the routing distribution shifts over training as the method stops spending updates on low-signal or already stable samples.

NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games

Tom\'a\v{s} Hole\v{c}ek, Viliam Lis\'y Model-based reinforcement learning (MBRL) has worked well in single-agent settings, but extending it to two-player zero-sum imperfect-information games (IIGs) runs into opponent-induced non-stationarity and identifiability barriers that, the authors argue, make centralized model learning a mathematical necessity. NashDreamer introduces a centralized Multi-Agent Recurrent State-Space Model (MARSSM) that separates environment dynamics from the effect of each player's strategy on their own observations, and it can use any policy gradient algorithm while inheriting that algorithm's convergence guarantees toward Nash equilibria under an idealized model. On four benchmark games, it substantially improves sample efficiency over model-free baselines early in training. A theoretical analysis of the optimization landscape also identifies a vulnerability of Dreamer-style algorithms to posterior collapse in stochastic environments, which the authors leave as an open challenge.

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

Matteo Merler, Giovanni Bonetta, Davide Zago, Rossella Cancelliere, Bernardo Magnini Vision-Language Models (VLMs) carry useful priors for interactive decision-making, but using them directly as policies is expensive, does not improve with experience, and can repeat systematic errors. SAGE (Selective Agent Guidance via Entropy) queries a VLM teacher only when a lightweight reinforcement learning (RL) learner is uncertain, executes the suggested action during training, and distills the guidance into the learner's policy, optionally weighting teacher actions by environment-derived advantages because the advice is not always reliable. Across sparse-reward visual reasoning and navigation tasks, the resulting policies act without any VLM calls at evaluation time and improve over unguided RL in several environments, including settings where the learned policy exceeds its VLM teacher. Guidance helped most when the VLM steered the agent toward high-reward trajectories and least when unguided exploration already succeeded or teacher actions produced uninformative experience.
2 more specialized papers

Robotics 11

IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

Rongze Tang, Jianjie Fang, Zhaolu Wang, Ziyou Wang, Xvyuan Liu, Haisheng Su et al. World models for embodied agents predict future frames from actions but generate physically implausible interactions, and the usual remedy injects external motion, geometry, or semantic representations that require auxiliary estimators or manual annotation. The diagnosis offered instead is a supervision-allocation mismatch in the globally averaged mean squared error denoising objective: abundant static content dominates the optimization signal, leaving the sparse dynamic regions that carry the interaction under-supervised. IMPACT uses cross-attention on manipulated-object tokens as an internal spatiotemporal prior, samples candidate regions from it, calibrates them with detached local prediction errors into an interaction map, and reweights the denoising loss accordingly, requiring no external representations and no inference-time modifications; across robot-arm and human-hand manipulation with several diffusion transformer backbones it improves interaction fidelity, physical plausibility, and visual quality over the MSE-trained baselines.

A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies

Ahmad Alfan Alfian Irfan, Nur Ahmad Khatim, Mansur Arief Driving policies deployed on embedded hardware are shrunk by pruning, distillation, and quantization, but the aggregate scores used to sign off on the compressed model may not reflect whether the car still drives safely around other road users. The authors train a belief-state policy with proximal policy optimization (PPO) in Gym-Duckietown, then push the extracted actor through the compression pipeline one stage at a time, evaluating in closed loop on five driving curricula after each stage. Structured pruning is the stage where driving capability first disappears; distillation partially repairs the pruned actor but is limited by its rehearsal data, and integer quantization then costs the curricula requiring the vehicle to stop and resume — whereas quantizing the unpruned actor preserves all five.

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

Wei Wang, Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin, Jinhao Mao et al. cross-listed Vision-language-action (VLA) models map observations and instructions to robot actions, but long-horizon tasks also require coordinating perception, planning, execution, progress verification, and recovery as the physical state changes. EmbodiedSkills treats each skill decision as an execution proposal whose prerequisites are checked before execution and whose outcome is verified afterward, connecting high-level skill selection, bounded low-level VLA execution, and post-action verification through a fixed executable-skill interface that also logs structured trajectories usable as supervision for individual components. Instantiated with Qwen3-VL and OpenPI/pi0.5, task-adapted low-level policies reach 86.20% average success across 50 RoboTwin 2.0 tasks and 97.40% across the four LIBERO suites, though the same approach achieves only 12.5% on four memory-dependent RMBench tasks.

Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching

Jaewoo Park, Minyoung Lee, Sukmin Seo, Moonbin Yim, Hyunwook Yoon, Dohoon Ryu et al. cross-listed How far multimodal large language models (MLLMs) extend from perceiving to acting is tested by dropping one directly into a drone's control loop with its entire action space declared solely in the prompt. DroneCATS-Agent makes the model a swappable component and the DroneCATS benchmark treats it as the independent variable across approaching a visible target, tracking a moving one, searching outside the initial view, and commanding a multi-drone fleet, with no fine-tuning or function-calling schemas and a roster scaling down to 2B parameters. Small open models often navigate into the success radius more reliably than frontier models yet lose the episode by declaring arrival prematurely or never at all, and multi-drone commanding widens the gap as small models blindly copy one coordinate across distinct views, making protocol discipline and correct termination, not navigation, the bottleneck.

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds

Clinton Enwerem, John S. Baras, Calin Belta cross-listed Imitation-learned manipulation policies are usually stress-tested against variation in scenes, objects, or instructions rather than in how fast the task is executed, leaving open how much of the expert's speed tolerance the learner inherits. The comparison uses ParcelStow, a contact-rich task where a robot acquires, reorients, and inserts a parcel; a scripted expert and an ACT (Action Chunking with Transformers) policy trained on its demonstrations both reach 100 percent success at nominal speed. At the fastest demonstrated speed the expert still succeeds 84 percent of the time while ACT falls to 53 percent, with 35 of its 47 failures being insertion misalignments, and across all policies and speeds none of the 414 acquisitions lacking force closure ever completed the task. Matching success at nominal speed therefore says nothing about whether temporal robustness survived imitation.

Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation

Haoyuan Deng, Haichao Liu, Wenkai Guo, Yuan Ling, Zaijia Yang, Yuanjiang Xue et al. cross-listed Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. Facet-0 unifies multimodal representation learning and reinforcement learning (RL) post-training around a joint action-wrench proposal: a causal wrench history is aligned with vision-language semantics and kinematic state, flow matching generates each action chunk together with the wrist-wrench profile it is expected to induce, and deployment rollouts train a distributional Action-Wrench Critic that separates motions with similar task progress but different contact outcomes, with phase-aware rewards and contact-selective credit focusing improvement on decisive interactions. A lightweight bounded actor reuses the frozen representation for on-robot adaptation to part-specific dynamics while RL stays defined over executable Cartesian actions. Trained on ManuFacet-1K, a 1,000-hour force-synchronized corpus spanning three embodiments and multiple manufacturing cells, the task-adapted system reaches 82% mean success on five sub-millimeter computer-assembly tasks versus 15% for the strongest baseline, with 0.5 mm placement accuracy and 50 ms command latency.
5 more specialized papers

Reasoning 8

RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving

Xiyuan Zhou, Zhuoqi Li, Xinlei Wang, Yirui He, Yuhao Wu, Yuheng Cheng et al. Rewriting benchmark problems is a common defence against data contamination in evaluating LLM mathematics, but current rewriting methods cannot guarantee that the new problem is well posed or that the new answer is right. RePro folds Lean-oriented neural automated theorem provers into the rewriting loop, regenerating both problems and answers with correctness backed by machine-checked Lean proofs. On GSM8K and MATH, the instances RePro retains hit 100% well-definedness, feasibility, and answer correctness where prior methods still emit invalid or wrongly answered items, and several models lose accuracy on the verified rewrites — a sign their scores partly reflect memorization of surface form and structure.

Dependency-Aware Chain-of-Thought Compression for Financial Reasoning

Wenjun Wu, Lei Fu, Kejian Tong, Tao Ning, Sichen Zhao Chain-of-thought prompting helps on hard reasoning but its long intermediate traces drive up inference cost enough to obstruct deployment in financial settings. The Hierarchical Semantic Distillation Network (HSDN) compresses those chains through semantic segmentation, dependency graph construction, dual-encoder importance scoring, constrained segment selection, and local boundary rewriting, using a frozen Qwen3 4B only for feature extraction and final answer generation so the compression stage itself stays structured and interpretable. On the AFAC2025 benchmark it reaches 91.0% accuracy at 68.4% compression, ahead of strong compression baselines on overall score and reasoning coherence.

Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs

Lu Cheng Inference-time search with language models tends to pile its budget onto a handful of structurally similar trajectories, a failure the authors name reasoning basin collapse. BASIN is a training-free selection rule that clusters reasoning states into basins and penalizes revisiting the same strategy, redirecting a fixed compute budget toward genuinely different paths; a quality-aware variant, QA-BASIN, keeps strong basins when pure diversification overshoots. Against Tree of Thoughts under matched budgets it gains up to 22 percentage points on Game of 24 and 6.7 points on MuSR, and a proposed redundancy gap statistic — how differently search concentrates on correct versus incorrect predictions — shifts from roughly zero under Tree of Thoughts to consistently positive.

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory

Xincheng Wei, Yifan Ding, Yoshua Li, Dongsheng Ma, Rongxiang Weng, Xunliang Cai et al. Self-play lets a model generate its own training questions, but without direction its solver plateaus: existing unguided signals such as difficulty, learnability, or diversity keep questions hard and varied without saying which reasoning weaknesses to attack, while guided methods import direction from human examples or document corpora outside the loop. DiagEvo derives that direction from the solver's own failure history — a diagnostician extracts recurring error causes into a hierarchical memory that groups them under skill nodes and marks each Active or Mastered by self-consistency on targeted questions, and the challenger uses those states and recurrence counts to trade off cause-targeted generation against free exploration, with double-confidence filtering keeping mid-difficulty questions only when the majority answer leads clearly. With a 4B diagnostician it beats every baseline in mean accuracy across all nine benchmarks for Qwen3-4B, Qwen3-8B, and OctoThinker-8B, reaching 72.3% mean accuracy on five mathematical reasoning benchmarks with Qwen3-8B, 4.5 points above R-Zero.

StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?

Yinghao Chen, Zixi Chen, Bingxiang He, Ziqing Qiao, Huan-ang Gao, Yinuo Xu et al. Self-evolution methods claim to let models learn autonomously from raw material, but there has been no direct way to measure how efficiently they convert study material into problem-solving ability. StudyBench is a controlled physics benchmark that splits evaluation into an Application Set of hard textbook problems, testing absorption, and a Transfer Set of olympiad-level problems, testing generalisation. Benchmarking representative self-evolution methods on three base models shows gains on the Application Set rarely carry over to the Transfer Set, and an ablation exposes a "Guidance Gap": the strongest method recovers only a small fraction of what the same material unlocks when simply supplied as in-context guidance. Every method also plateaus well before its compute budget is exhausted, which the authors read as evidence that the bottleneck is methodological rather than a shortage of data or compute.

Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs

Zhaoliang Chen, Jie Fu Chain-of-thought reasoning commits each step as text, so errors propagate and training requires traces to imitate; reasoning in a model's continuous representation space avoids these constraints but leaves open how the latent states should be computed. Latent Recurrent Thoughts (LRT) keeps a large language model (LLM) frozen as the decoder while a task-dedicated proposer supplies base latents and a tiny recurrent reasoner refines them over many steps through bounded residual corrections, decoupling depth of computation from model size. On Countdown-4, Sudoku, HumanEval, MBPP, and StrategyQA, LRT substantially outperforms prior frozen-decoder continuous-space reasoning methods under an identical decoder, prompt, data, and training budget, and it beats non-thinking-mode chain-of-thought prompting on the same backbone at a small fraction of its inference compute.

Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning

Mariia Drozdova, Aidan Sirbu, Pietro Miotti, Robert Obryk, Mayalen Etcheverry, Eyvind Niklasson et al. Diffusion denoisers and recursive reasoners both iterate but differ in what they carry between steps; adding a persistent hidden state to a denoiser and removing its timestep conditioning leaves a single shared update that can run to arbitrary depth. The resulting anytime solver keeps improving well past the rollout lengths and backpropagation window used in training, reaching 99.90% exact solve on Sudoku-Extreme and 98.93% on Maze-Unique. Progressive denoising turns out to be unnecessary at inference: holding corruption at maximum by replacing every non-clue variable with fresh Gaussian noise each step still converges to stable correct answers, with no parallel rollouts, candidate selection, or external verifier. Ordered annealed corruption remains essential during training, suggesting diffusion's contribution here is a denoising curriculum rather than a sampling procedure.
1 more specialized paper