Monday, August 17, 2026

284 papers cs.AI · cs.LG · cs.CL ← 2026-08-142026-08-19 →

Jul Aug Sep

Highlights

MobileMem: Learning from a Year of Mobile Experiences

Highlight HF pick · 15▲Agents Xinle Deng, Yida Xue, Xiangyuan Ru, Haoming Xu, Shuofei Qiao, Mengru Wang et al. Personal assistants that accumulate knowledge about a user over months need benchmarks built from heterogeneous, multimodal, evolving personal data, which existing long-term memory evaluations do not provide. MobileMem supplies both a benchmark and framework for on-device memory, using a knowledge-grounded synthesis pipeline to turn a year-scale collection of real user-app sessions into coherent, temporally consistent long-horizon trajectories. It offers matched text-only and multimodal settings that probe multi-hop and temporal reasoning, updating superseded knowledge, and inferring preferences the user never stated, framing memory as accumulated experience rather than a retrieval index over isolated facts.

Persistent mobile assistants need memory spanning a year of heterogeneous, cross-app, multimodal activity, but existing long-term-memory benchmarks assume cloud storage and clean dialogue histories. MobileMem synthesizes year-scale interaction trajectories from real smartphone usage metadata and interviewed personas, then evaluates nine memory systems end-to-end on multi-hop, temporal, preference, and unanswerable questions.

  • The KEME synthesis engine alternates top-down temporal planning — recursively expanding a persona and time horizon into an event graph, then into sessions — with bottom-up experience evolution that revises unexpanded future events and updates persona attributes with message-level evidence links, using real OPPO app-usage metadata as immutable "knowledge anchors"; the text track covers seven app templates while MobileMem-Omni adds screenshots, synthetic portraits, and bilingual dialogue in trajectories exceeding 2M tokens.
  • Systems that preserve raw conversational detail win: A-MEM scores 79.68 overall and HippoRAG2 78.85 with a GPT-4.1-mini backbone (HippoRAG2 reaching 80.06 on GPT-5.4-mini), while a 128k-token long-context baseline reaches only 54.51 and compress-and-overwrite designs like Mem0 (35.63) and LangMem (24.79) collapse.
  • Memory construction cost varies by roughly 4x for similar accuracy — A-MEM spends 5.46M tokens per trajectory on GPT-4.1-mini and 11.17M on GPT-5.4-mini (which extracts far more keywords per memory unit), whereas HippoRAG2 matches or beats it at about 2.8M.
  • Temporal reasoning is the consistent weak spot, topping out at 72.04 against 86.30 on single-hop, and adversarial unanswerable questions invert the whole ranking: LangMem, last overall, leads at 76.42 because stronger retrievers surface weakly relevant distractors that convince the model an answer exists.
  • The main caveats are that trajectories are LLM-synthesized by GPT-5.2 from only two volunteers in the text track and eight real plus eight virtual personas in Omni, global consistency is checked by manual sampling rather than any automatic metric, only about 44% of fine-grained profile fields ever surface in the generated data, and — despite the on-device framing — every evaluated system runs cloud-hosted backbones with cost reported purely in tokens, never in storage, latency, or power.

Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning

Highlight Theory Yongchao Huang A hidden Markov model (HMM) does three things — infer a belief over hidden state, propagate it through a transition, and emit back into observation space — and the argument here is that time-indexed Predictive Information Bottleneck VJEPA (PIB-VJEPA) has exactly that structure, with the stochastic context encoder acting as an amortized filter, the probabilistic predictor as latent dynamics, and the decoder or inverse target encoder as the emission. Four progressively stronger levels of correspondence are defined along with sufficient conditions for exact sequence-level HMM equivalence, and MCJEPA makes the link concrete by replacing the latent predictor with a learned transition matrix whose powers guarantee exact multi-horizon Chapman-Kolmogorov consistency in the finite time-homogeneous case. Deterministic temporal JEPA falls out as a degenerate Dirac-kernel case, and controlled experiments confirm transition composition, the filtering interpretation, and predictive Markovization on a synthetic process with known structure.

Predictive-coding architectures like JEPA are usually described in terms of representation matching rather than probabilistic inference, leaving unclear what generative model, if any, they correspond to. The claim here is that a fully stochastic temporal JEPA — specifically PIB-VJEPA, where the current representation, the future target, and the transition are all distributions — instantiates the encode–transition–emit pipeline of a hidden Markov model, with the context encoder acting as an amortized filtering distribution, the predictor as the latent transition, and a decoder or inverse target encoder as the emission.

  • The correspondence is graded rather than binary, split into four levels — shared computational roles, an emission-complete latent state, sequence-level HMM equivalence, and model-and-objective equivalence — with a theorem giving four sufficient conditions for the third level: first-order Markov latent dynamics, a valid state-to-observation conditional, transition-consistent latent marginals, and a history encoder that coincides with the induced filtering posterior.
  • MCJEPA (Markov-Chain JEPA) makes the idea concrete by replacing the neural predictor with a learned row-stochastic transition matrix A over K categorical states, so an h-step forecast is just q_t A^h and matrix powers give exact Chapman–Kolmogorov path consistency — every decomposition of the same horizon yields an identical predictive distribution, which is what makes hierarchical planning over composed short- and long-horizon steps well-defined.
  • Because the encoder, target encoder, and transition are trained jointly against a KL latent-matching loss, they can agree through degenerate solutions, so training adds an occupancy term KL(q̄_B || Unif(K)) against single-state collapse plus a per-sample entropy penalty against uniform-assignment collapse — two regularizers whose weights must be balanced, since too much occupancy pressure forces artificial state usage and too much entropy pressure causes premature hard assignments.
  • The construction generalizes cleanly along both axes of latent Markov dynamics: side-information-conditioned matrices A_φ(ξ_t) for action- or time-dependent transitions, continuous-state Gaussian or flow-based kernels, continuous-time generators via exp(Δt Q_φ) for irregular sampling, and deterministic temporal JEPA recovered as the degenerate Dirac-kernel limit as the transition covariance goes to zero.
  • The honest limits are substantial: standard JEPA training optimizes latent target matching plus information-bottleneck regularization, not observation-sequence likelihood, so the stated conditions are sufficient but not necessary and are not guaranteed by ordinary JEPA training; compressive target encoders are generally not invertible, the implicit Bayes-rule emission p(x|z) ∝ p_data(x) q_θ(z|x) is a static correspondence that need not be tractable for generation; and the empirical support comes from 4 controlled experiments on synthetic processes rather than any real video or large-scale benchmark.

Second Thought: Reasoning in Parallel as LLM Agents Act and Observe

Highlight HF pick · 5▲Agents Zhensu Sun, Chengran Yang, Yunbo Lyu, Jieke Shi, David Lo Agents in the ReAct paradigm think only during the Thought phase and sit idle while an action is serialized and the environment responds, leaving a recurring reasoning idle window unused. Second Thought is a training-free inference framework that forks four auxiliary reasoning branches the moment a Thought phase ends, decodes them concurrently with the main loop, and merges their output back when the observation arrives, moving the extra reasoning off the sequential critical path. Across three agentic benchmarks and three reasoning models it lowered average turn count in all nine model-benchmark pairs and cut main-thread decoding by up to 43% with Pass@1 unchanged in seven pairs; against a compute-matched control that spends the same budget on the main thread, it achieved strictly higher Pass@1 with 1.3 to 3.2 times less sequential decoding.

ReAct agents only reason during the Thought phase, leaving the Action–Observation interval — the wait for a tool to execute and return — as a recurring window in which no reasoning happens at all. Second Thought is a training-free inference framework that forks four auxiliary reasoning branches the moment a Thought concludes, decodes them concurrently with the main loop, and appends whatever they produced to the observation message so the next turn starts with the extra deliberation already in hand.

  • The four branches each take a complementary angle on the same live trajectory — Check audits the just-finalized plan for unverified assumptions, Recall resurfaces constraints fading from the context window, Rehearse pre-computes conditional next steps, and Alternative drafts fallback strategies — so merging is plain concatenation rather than voting or ranking, and all four share the main thread's prompt prefix KV cache.
  • Because the window closes unpredictably when the observation arrives, every branch emits a stream of atomic thoughts: self-contained units of ≤25 words wrapped in <thought> tags with no cross-references, so cancelling mid-generation discards at most the unit in flight, and each buffer is truncated at its last closed tag and capped at 5 thoughts per dimension.
  • Across SWE-Bench Pro, Terminal-Bench 2.1, and τ³-bench with DeepSeek-V4-Flash, Qwen3.6-Plus, and MiniMax-M3, turn count falls in all nine model–benchmark pairs and main-thread decoding in six by up to 43% (roughly 20% on average there), while Pass@1 is statistically unchanged in seven pairs and significantly higher in two (+12.4 and +10.2 points, both on Terminal-Bench 2.1); a paired wall-clock replay of 50 SWE-Bench Pro instances turns this into 10.9% lower median per-task latency (256.9s to 229.0s), with branch contention costing only 2.8s against 27.9s saved.
  • Against s1 budget forcing — a compute-matched control that spends the same extra reasoning tokens on the main thread's own thought — Second Thought reaches strictly higher Pass@1 with 1.3× to 3.2× less sequential decoding in all four settings where the control is applicable, and a replay ablation shows harvested thoughts substitute for work the agent would otherwise do inline (next-turn reasoning grows from 196.2 to 316.5 tokens when they are removed).
  • The main costs are monetary and structural: four branches raise per-task API spend by 66.4% to 181.5% (almost entirely cached prefix reads, reducible to 16.3–35.5% by keeping only the Alternative branch), the default configuration is window-starved — letting branches run to completion on the critical path reaches 56.7% Pass@1 versus 52.0%, meaning the truncated version captures only 41% of the attainable gain — and on τ³-bench, where windows are short and failures stem from retrieval and policy adherence rather than planning, gains top out at +3.1 points.

From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models

Highlight Large Language Models Pranav Kumar Kaliaperumal A retrospective covering October 2018 to July 2026 tracks how language models went from BERT to agents that solve competition math and write software, reporting that the ability to resolve real coding issues improved nearly sixfold per year since late 2024 while cost per unit of capability collapsed, with budget-tier models matching flagship capability at one to six dollars per million tokens. It also documents fragmentation of the frontier into task-targeted models, with different systems leading on frontend coding, repository-level coding, and terminal tasks. A companion grade-school math experiment on Qwen2.5 shows basic decoding solving 58 of 100 problems against up to 79 with advanced sampling, and a confidence ranker placing 47 correct answers in its top 50, with all materials released publicly.

Eight years of language-model progress are measured against public benchmarks, API prices, and one reproducible small-model experiment, rather than against impressions. The picture that emerges is a capability curve still compounding fast, a cost curve collapsing at a similar rate, and a frontier that has fragmented into task-specific leaders — so the practical unit of deployment becomes a routing policy rather than a single model.

  • A log-linear fit to SWE-bench Verified release scores from October 2024 to July 2026 puts annual growth in the odds of resolving a real GitHub issue at ~5.8× (R²=0.78, n=14), carrying the independent frontier to 96–97% and implying the benchmark saturates within about a year, the same lifecycle GLUE and MMLU already completed.
  • Input-token prices fell ~60× from GPT-3's $60/M (2020) to GPT-5.6 Luna's $1/M (2026), and in OpenAI's own launch table Luna wins 7 of 10 agentic and professional-work benchmarks against the eleven-week-old GPT-5.5 flagship at one-fifth the price — while trailing sharply on long-context recall (41.3% vs 81.5% on MRCR) and hardest academic reasoning.
  • The mid tier is squeezed out: a Pareto analysis of fifteen GPT-5.6 effort settings finds every Terra configuration dominated by a cheaper-or-smarter Luna or Sol setting, and effort itself has an interior optimum — Opus 5 peaks on Frontier-Bench at xhigh (44.4%) and drops to 43.3% at max.
  • No model leads everywhere — Opus 5 tops frontend preference (1,712 Elo) and ARC-AGI-3 (30.2%, ~4× the prior 7.8% record), Fable 5 leads SWE-bench Pro at 80.0% against Sol's 64.6%, and Sol leads terminal work — yet a two-model router pairing Sol with Fable 5 recovers the entire +2.4-point gain of a six-model oracle on the fourteen-benchmark suite.
  • The inference-time experiment is honest about its own weakness: a frozen Qwen2.5-1.5B configuration on 100 locked GSM8K items scores 58/100 greedy versus 62/100 for four-sample plurality (paired McNemar p=0.481, not significant) against a 79/100 any-sample oracle, so selection — not sampling — is the bottleneck; a post-hoc confidence model reaches 0.833 AUC and 47/50 correct in its top-ranked half, but it was designed after labels were visible and cross-validated on those same items.
  • Provenance is the standing caveat throughout: mid-2026 numbers mix vendor claims with independent harnesses (GPT-5.5 scores 85.1% vendor-reported versus 82.6% reconstructed), the oracle router cheats by knowing benchmark identity, and targeted gains may not transfer — Opus 5's ARC-AGI-3 record collapses to a statistical tie on the held-out Witness puzzles.

Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation

Highlight Safety & Alignment Ajay Pravin Mahale (Hochschule Trier) The EU AI Act obliges providers of high-risk systems to document how their systems reach decisions, and circuit discovery in mechanistic interpretability is the natural source of such evidence — provided two competent analysts using the same tool with different defensible settings would file the same thing. A pre-registered grid crossed seven analytic axes, each level drawn from a published implementation, over GPT-2 small on the indirect object identification task, mapping every discovered circuit through a deterministic claim map into a structured Annex IV statement. Across 15,840 specifications, 7,561 of which produced a claim, the derived statement flips across 73.2% of specification pairs and the most common claim covers only 41.1% of the space; standardizing the evaluation metric leaves the flip rate at 59.4%, and dropping circuit size from the claim entirely still leaves 27.1%. The underlying circuits are near-disjoint (median pairwise Jaccard overlap 4%) and functionally uncorrelated (Cohen's kappa 0.015), so this is not one mechanism described in different words; the study covers one model and one task.

Circuit discovery is the most mature instrument mechanistic interpretability could offer to satisfy EU AI Act Annex IV documentation duties, but its output is only useful as evidence if two competent analysts using the same tool on the same system file compatible claims. A pre-registered multiverse over seven defensible analytic axes, with each discovered circuit mapped through deterministic code into a structured regulatory statement, finds that the filed claim flips across 73.2% of specification pairs.

  • The design crosses seven axes — discovery objective, ablation operator, corruption distribution, evaluation metric, size threshold, prompt variant, seed and granularity — with every level taken from a published implementation, yielding 15,840 specifications on GPT-2 small and the indirect object identification task, of which 7,561 survived the metric-relative threshold rule and produced a claim.
  • The modal claim commands only 41.1% of the space, failing a proposed filability criterion (modal share ≥ 1 − α) at every tolerance a conformity body would plausibly accept, and the top two claims contradict rather than merely differ: 41.1% attribute the behaviour to early layers and 28.5% to late layers.
  • No single control rescues the filing — standardising the evaluation metric, the most influential axis, leaves the flip rate at 59.4%, while removing circuit size from the claim entirely and holding size fixed still leaves 27.1% (95% CI 0.255 to 0.286).
  • The instability is structural and functional, not verbal: median pairwise Jaccard overlap between circuits is 4% and per-example functional agreement is uncorrelated at Cohen's kappa 0.015, which directly rebuts the "phantom specialisation" objection that discovery merely samples behaviourally equivalent subgraphs.
  • Caveats are substantial and the authors foreground them — one model and one task (a Pythia-160m replication is pre-registered but unrun), a 52.3% discard rate skewed heavily by metric, one of the library's seven documented discovery objectives that crashes on its own canonical task, and a pooled-versus-within-size reversal showing discovered circuits are in fact more stable than a size-matched random null (0.2746 against 0.4230) once size is held constant.

Scaling Domain Data Repetition in LLM Pretraining

Highlight HF pick · 3▲Large Language Models Jingwei Li, Xinran Gu, Rui Dai, Xintong Hao, Chengyin Xu, Yan Wu et al. As models grow, token budgets grow with them to keep a sensible tokens-per-parameter ratio, but scarce high-quality domain data cannot scale the same way and gets diluted in the mixture; repeating it counteracts dilution at the risk of overfitting. Sweeping this trade-off under practical scaling, where token budget grows with model size, yields two findings: at fixed tokens-per-parameter the optimal repetition count mildly increases with model size, and across domains the optimal count correlates strongly and negatively with a domain's final validation loss, while the amount of unique domain data barely matters. The practical implication is that repetition counts tuned on small proxy models at the same tokens-per-parameter ratio carry over to larger runs.

Practical LLM scaling grows the training-token budget in step with model size, but high-quality domain data (code, math, curated text) cannot be scaled the same way, so its share of the mixture is progressively diluted. Repeating that scarce data counteracts the dilution at the cost of overfitting risk, and this work charts where the optimum sits as a function of model size, domain, and corpus size.

  • The setup sweeps repetition counts for a fixed domain while holding the tokens-per-parameter ratio (TPP) constant and scaling the token budget proportionally with parameters, using final per-domain validation loss as the criterion for the optimal number of passes.
  • Contrary to the usual intuition that larger models memorize faster and should therefore see repeated data less often, the optimal repetition count mildly increases with model size at fixed TPP, meaning repeat budgets tuned at small scale are a floor rather than a ceiling.
  • Across domains, the optimal repetition count is strongly negatively correlated with a domain's final validation loss — domains the model already fits well can absorb more passes, while high-loss domains saturate sooner.
  • The amount of unique domain data is only weakly related to the optimal repetition count, undercutting the common heuristic of setting epochs from corpus size alone.
  • The headline practical claim — that repetition counts tuned on small proxy models at matched TPP estimate the right setting for large models — is supported by trend direction rather than by reported absolute numbers or correlation coefficients here, and rests on validation loss rather than downstream benchmark quality, so how far the mild upward trend extrapolates beyond the tested size range remains open.

Forecast Collapse in Time-Series Foundation Models

Highlight HF pick · 6▲Applications Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen et al. Forecasting hourly returns for 1,000 US equities makes time-series models emit nearly flat predictions that rank stocks poorly, a failure the authors name forecast collapse and that mostly vanishes when the target is trading volume instead. Sweeping time-series foundation models (TSFMs), twelve deep forecasters, and 97 benchmark configurations traces the effect to target predictability plus per-series training objectives that never identify cross-series structure, exposing a calibration-versus-ranking tradeoff: squared error flattens forecasts, while optimizing cross-sectional correlation directly can inflate amplitude by over an order of magnitude. Their CalibRank objective balances the two and nearly triples cross-sectional correlation on Finance1K while keeping amplitude near the target, improving correlation on every model tested. The broader point is that per-series evaluation metrics can hide the cross-series failures that downstream decisions depend on.

Forecasting hourly returns for 1,000 US equities produces predictions that are nearly flat and rank stocks poorly by cross-sectional correlation — a failure the authors call forecast collapse — even though the same models forecasting trading volume in the same setting behave normally. The core claim is that collapse tracks target predictability, and that the standard squared-error objective is structurally unable to deliver both calibrated amplitude and usable cross-sectional ranking.

  • The phenomenon is characterized across time-series foundation models, twelve deep-learning forecasters, and 97 public benchmark configurations, isolating two distinct causes: low predictability caps the amplitude any calibrated point forecast can take, while per-series training objectives never identify cross-series structure at all.
  • This yields a calibration-ranking tradeoff — minimizing squared error drives forecasts toward flatness, whereas directly optimizing cross-sectional correlation recovers ranking but inflates forecast amplitude by more than an order of magnitude.
  • CalibRank, the proposed objective, mixes the two terms to balance calibration against ranking, and on Finance1K it nearly triples cross-sectional correlation while holding forecast amplitude close to the target, improving correlation on every model tested.
  • The volume-versus-returns contrast is the cleanest evidence that this is about signal-to-noise rather than architecture, since identical models and data pipelines collapse on one target and not the other.
  • The broader methodological point is that conventional per-series metrics such as MSE or MAE can look healthy while the cross-series structure that downstream ranking decisions actually depend on has already degenerated.

Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

Highlight HF pick · 24▲Large Language Models Kai Chen, Jifeng Ding, Ning Ding, Jiaye Ge, Lixin Gu, Yicheng Gu et al. Mobius-v0 splits a transformer's usual entanglement of knowledge and computation into a globally shared Memory built from feed-forward layers that stores knowledge vectors, and multiple Reasoners built from self-attention that repeatedly query that memory using hidden states as both cache and carrier. Because knowledge lives in one shared store while reasoning is applied iteratively, the architecture compresses knowledge more tightly and reuses reasoning operators. A 7B model trained from scratch matched a 7B transformer baseline's downstream scores using 62.6% of the baseline's training data, and Intern-S2-Mobius, continually pretrained from Qwen3.5-35B, matched downstream quality while delivering nearly 4x end-to-end inference speedup.

Mobius-v0 breaks the Transformer's layer-wise binding between feed-forward knowledge storage and attention-based reasoning, replacing per-layer FFNs with a single globally shared memory that every attention "reasoner" can query. The claim is that this decoupling buys both better knowledge compression and a shorter path to an answer, since latents can be refined against the full knowledge pool in a few layers rather than traversing every layer per token.

  • Sharing one oversized key-value memory across layers gives the model an implicit backward residual connection — deep reasoners can reach shallow-layer knowledge, not just the reverse — plus dynamic latent reasoning in which hidden states iterate against the repository and decode several tokens at once; at scale the shared FFN is partitioned MoE-style with sparse activation to keep the cost tolerable.
  • Trained from scratch as a 7B-A1B MoE on 1TB tokens, Mobius reaches the Transformer baseline's final MMLU score using only 62.6% of the data, which the authors attribute to less cross-layer redundancy in how knowledge is stored.
  • Intern-S2-Mobius-35B, continually pre-trained from Qwen3.5-35B-A3B on 1TB tokens plus SFT and RL, averages 67.88 vs 65.05 on general benchmarks (AIME 2026 95.31, HMMT 2026 85.51) and 52.14 vs 18.20 on scientific ones like Mol-Instructions and MolecularIQ, while delivering roughly 4× end-to-end inference speedup.
  • The speedup comes almost entirely from shorter chains of thought rather than faster forward passes — on an MMLU-Pro linear algebra item the model reaches the same correct answer in 516 tokens against 2,364, skipping the baseline's 1,147 tokens of repeated derivation and checking.
  • Weaknesses are real and acknowledged: the converted model loses to its own base on UGD hard (73.02 vs 78.02) and HLE (19.11 vs 22.40), the mechanisms behind both the data efficiency and the shorter reasoning traces are hypotheses rather than established results, expert routing still carries a block-diagonal prior inherited from the source checkpoint, and the compositional-generalization evidence in the appendix comes from an unreleased newer architecture.

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

Highlight HF pick · 43▲Vision Shuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen, Zihan Mei, Manting Li et al. Modern video generators can produce convincing footage of wars, disasters, and other emergencies, but existing detection benchmarks say little about how detectors behave on this kind of content or after it spreads online. RA-Bench pairs 1,830 real videos across 10 social-risk categories with 16,056 clips generated from four open-source and five closed-source generators, then evaluates seven traditional detectors, ten zero-shot multimodal models, and two multimodal large language models fine-tuned for the task. None of the three detector families generalizes consistently across the benchmark, generation quality and conditioning affect each family differently while source-level patterns stay stable across sampling seeds, and the videos that fool human judges are also the hardest for detectors — with social dissemination (re-encoding and sharing) making detection harder still.

Video generators can now fabricate convincing footage of wars, disasters, and public emergencies, but existing detection benchmarks are built on generic web video and say little about whether detectors hold up on socially consequential events. RA-Bench closes that gap by anchoring 1,830 real crisis clips across 10 social-risk categories and conditioning nine image-to-video generators on each clip's first frame, yielding 17,886 videos used to stress-test three detector families, generation properties, and survival through social sharing.

  • Each real anchor is captioned with Gemini-3.1-Pro-Preview, and that caption plus the anchor's first frame is fed to four open-source and five closed-source generators, so every fake shares scene semantics and a genuine opening frame with its real counterpart — mimicking the realistic attack where a true crisis photo seeds a fabricated video.
  • Detection collapses out of domain: the seven traditional detectors fall from published AUCs of 67.6–98.6% to source-level means of 43.9–57.3%, with 26 of 63 detector–generator pairs scoring below chance, and the public ranking barely transfers (Spearman 0.26), so no single detector leads across generators.
  • MLLM detectors fail differently — Gemini-3.1-Pro-Preview tops the zero-shot models at only about 63% BAcc while smaller Qwen3.5 variants flip from 19.7% to 100% fake-recall purely on prompt format, and fine-tuned Skyra turns out to key on a timestamp artifact: swapping absolute timestamps for frame indices over identical frames drops it from 68.5% to 54.4–54.9% BAcc, while BusterX++ calls almost everything real (FakeR 4.1–9.1%).
  • The videos that fool people are the same ones that fool machines: 20 reviewers flagged 68.6% of open-source clips but only 52.9% of closed-source ones (Seedance2.0 40.7%, Kling 45.1%), and on the 633-clip RA-Bench-HumanProof subset that all five reviewers called real, Gemini reaches just 54.7/54.5% BAcc and traditional detectors average 47.5% AUC — worse than a coin flip.
  • Real-world circulation is the hardest blow: the full RA-Bench-LastMile dissemination simulation cuts mean fake-recall across five fine-tuned configurations from 46.0% to 1.4%, and more real-image conditioning (T2V → first-frame → first+last-frame I2V) steadily suppresses fine-tuned MLLM FakeR from 70.5% to 42.7% to 28.3%, though the study reports associations within one benchmark rather than causal effects and closed-source coverage is uneven because provider safety filters rejected some prompts.

Marionette: Predicting World States, Rendering Geometry, Painting Appearance

Highlight HF pick · 6▲Vision Zian Meng, Zhen Li, Chuanhao Li, Qiang Li, Kaipeng Zhang Interactive game world models usually autoregress pixels or latents directly, which forces pose, geometry, and occlusion to be tracked implicitly and lets errors compound over long horizons. Marionette splits the job three ways: a two-stage autoregressive dynamics model predicts an explicit 276-dimensional 3D world state of articulated skeletons, metric root trajectories, and rotations; a zero-parameter graphics bridge turns that state into pose-control videos with world-space geometry and occlusion computed in closed form; and a control-conditioned video diffusion model paints photorealistic frames. Because behavior lives in the explicit state, it can be corrected there — left unconstrained, two generated characters drifted 21.2 m apart against roughly 5 m in recordings with a third of frames showing ground penetration, while adding just a terrain collider and a separation cap to the state cut penetration by 66% with no change to the observation model. Forcing a mismatched action stream shifted root-aligned joint error by 31% across 48 held-out segments, and routing appearance through the predicted state cost little visual fidelity (FVD 831 versus 799 for recorded poses).

Interactive game world models usually autoregress pixels or latents, leaving pose, geometry, and occlusion to be maintained implicitly by the same generative sequence, so errors compound over long rollouts. Marionette instead predicts an explicit 276-dimensional articulated world state, hands exact geometry to a zero-parameter renderer, and leaves only appearance to a video-diffusion model.

  • Dynamics are split into a compact decision model (ActionGPT, ~2.5M parameters, one discrete action token per frame per entity) and a larger animation model (PoseGPT, ~25M parameters, producing the 258-D body pose), after which a deterministic bridge integrates roots, places joints in metric world space, and rasterizes a three-channel geometry buffer that Wan2.2-Fun-5B renders into photorealistic RGB.
  • Because the action is an explicit token, control is applied by overwriting it: feeding a temporally shuffled action stream degrades root-aligned joint error by 31% (0.272 m to 0.357 m) across 48 held-out segments, with the correct forced stream matching free-run quality at 0.272 m against 0.281 m.
  • Routing every frame through the predicted state costs no measurable fidelity, with an FVD of 831 against 799 for recorded pose and 975 for an end-to-end pixel-autoregressive baseline trained on the same footage, though the bootstrap intervals for the baseline comparison overlap.
  • Long-horizon failures live in the state and can be repaired there without touching the observation model: a free rollout drifts the two characters 21.2 m apart (recorded sessions stay near 5 m) with a third of frames showing ground penetration, while a terrain collider cuts the collision-frame ratio from 0.337 to 0.114 (66%) and a 6 m separation cap holds the pair at 5.1 m for 7% more foot-skate.
  • The main limits are scope and drift: the dynamics model is trained on a single monster type's 173-action vocabulary while the appearance model spans 27 monsters, appearance identity decays across diffusion chunks since only a first frame anchors it, and a reported negative result shows a differentiable penetration penalty was gamed by stretching the skeleton (bone-length error rising from ~3% to ~13%), which is why constraints are imposed as inputs and post-hoc projections instead.

Applications 74

BCIJelly: An integrated ecosystem for brain-computer interface research

Liyuan Han, Xinrui Yang, Tianyu Zheng, Qizhi Yang, Yitao Qin, Liang Chen et al. cross-listed Brain-computer interface (BCI) research is slowed by incompatible data formats, one-off decoder implementations and hardware-specific deployment toolchains. BCIJelly bundles 18 curated datasets, 15 benchmark decoders, 80 reusable algorithm modules, an automated architecture search that builds task-specific decoders without manual design, and a toChip compiler targeting neuromorphic hardware, all in one Python framework with a graphical interface for non-programmers. The architecture search additionally runs in a closed-loop mode steered by a language model that reads task specifications, module descriptions and search history, and the stack is validated across five paradigms — motor, visual, speech, emotion and auditory — on recordings from humans, macaques and mice in single-task, multitask and cross-species settings.

Unknown Unknowns: Model Misspecification in Machine Learning for Physics

Juan Cruz-Martinez, Carolina Cuesta-Lazaro, Alexander Held, Michael Kagan cross-listed Inverse problems in particle physics and astronomy increasingly rely on models trained on simulation and deployed on real data, which raises the question of whether those models are wrong in unanticipated ways rather than merely whether they fit. The discussion frames misspecification as double-edged in physics — sometimes it is the discovery signal, sometimes a nuisance to absorb — and argues that a robust analysis absorbs the misspecifications one does not care about while preserving sensitivity to the ones one does. It surveys diagnostics for detecting misspecification and mitigation strategies, concluding that no single diagnostic can certify correct specification, so detection and mitigation must run as an iterative loop over a battery of complementary checks.

Building AI-Intensive Software with AI: Early Results and a Cautionary Tale on Measuring Development Cost

Victor Barros de Miranda Neves, Kiev Santos da Gama, Vinicius Cardoso Garcia cross-listed Reports on what AI-assisted software development actually costs are rare, and the few that exist are easy to get wrong in ways the final ratio never reveals. A six-person student team built a conversational onboarding assistant — retrieval-augmented code chat, guided tours, dependency graphs, technical-debt analysis — over one academic term with pervasive AI assistance, instrumented by a three-layer cost model tracking real AI spend, self-reported human effort, and a human-only counterfactual. The initially reported 19.4x cost advantage turned out to contain two independent mistakes, inferring per-token cost under a flat-rate subscription and pricing the counterfactual at the wrong regional labor rates, which together inflated the figure by roughly 2x; the corrected ratio is about 9.9x. The authors present the correction itself as the finding, since both errors are invisible in the headline number and plausibly common in similar reports.

Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu cross-listed Work on clinical citations from chatbots has focused on fabricated references, leaving open whether the studies they do cite are the ones expert reviewers would pick. Three general-purpose assistants — Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5 — were prompted with clinical questions adapted from 20 Cochrane Database of Systematic Reviews questions under patient, clinician, and researcher personas, producing 720 responses benchmarked against each review's included and excluded study sets. Responses recovered 39.2% of the studies Cochrane reviewers included on average while citing only 5.0% of excluded ones, with recall varying sharply by model (63.1% for ChatGPT versus 17.3% for Gemini) and modestly by user role. After controlling for publication year, citation rate, and open-access status, trial sample size was the only independent predictor of retrieval (odds ratio 1.80 per log unit), indicating a systematic pull toward larger trials.

PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization

Ruogu Chen, Jie Han Macro placement drives a chip's final performance, power, and area, yet placers optimize half-perimeter wirelength, a proxy that recent benchmarking found nearly uncorrelated with post-route timing — all six AI placers evaluated on ChiPBench were worse than the hierarchical baseline. A label-fidelity study across ten circuits and four flow stages identifies post-global-routing metrics as the best trade-off between fidelity to final timing and label cost, so PPAPlace trains a dual-stream surrogate — graph attention over the netlist plus spatial convolution over the placement grid — on those labels and backpropagates predicted worst and total negative slack gradients to cell coordinates. Used both as a co-objective inside an analytical placer and as a projected-gradient refinement of macro positions, it delivers 22% better worst negative slack and 51% better total negative slack than the hierarchical baseline on five held-out circuits without retraining.

Recent Advances in Deep Learning-Based Drug-Target Binding Affinity Prediction

Jafin Khan, Md Hossain Shuvo Drug-target binding affinity prediction is a core sub-problem of computational drug discovery where reported benchmark numbers have outrun real progress. Representative recent deep learning approaches are reviewed alongside seven widely used benchmark datasets and the standard evaluation metrics, with attention to architecture and molecular representation choices. The analysis finds that strong headline scores often reflect dataset bias and narrow evaluation protocols, and that most methods degrade substantially in cold-start settings involving unseen drugs or targets, pointing toward better dataset design, standardized evaluation, and multimodal representations.

HI-MeshGraphNets: Efficient and Accurate Mesh-based Physics Learning with Hierarchical Multi-scale Graph Neural Networks

SiHun Lee, Dong-Hyuk Park, Taesoo Bang, Seung-Hoon Kang Graph neural network surrogates for mesh-based simulation pass messages one hop per layer, so capturing long-range interactions on high-fidelity meshes demands deep processors that are slow, memory-hungry, and prone to over-smoothing. HI-MGN replaces the flat processor with a hierarchical one that coarsens the graph using farthest-point sampling and Voronoi partitioning while preserving the original mesh topology, letting information travel far in few layers, and reconstructs fine-resolution features with a learned graph interpolation network. On three structural and fluid benchmarks it delivers better accuracy than both MeshGraphNets and the Bi-Stride Multi-Scale GNN while cutting training time and peak memory.

Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Developmen

Jiada Li, Xuesong Ye, Olamide Olowoniyi cross-listed Team-level evidence for how AI coding assistants have changed open-source development practice is still thin, so this study tracks seven engineering metrics across all merged pull requests in two fast-moving AI infrastructure repositories, vLLM (18,290 pull requests) and SGLang (14,938), segmented into four eras. Throughput rose 21-fold and 17.9-fold respectively, and bot-authored pull requests accounted for under 0.2% of that growth, meaning the surge was overwhelmingly human-authored. Median cycle times in the latest era fell to roughly a day or less (with 90th-percentile tails above two weeks), comment density rose about four-fold with bots contributing an estimated 15-20% of that rise, contributor counts grew steadily, and pull request size stayed flat.

CipherSight: Robust Website Fingerprinting via Record-Resource Semantic Supervision under Distribution Shifts

Runhan Song, Qiqi Liu, Chuanzhou Pan, Zhenquan Ding, Youquan Xian, Chongru Fan et al. cross-listed Website fingerprinting infers which site a user visited from metadata leaking out of encrypted HTTPS traffic, but deployed classifiers degrade badly when traffic shifts across time and geography or when unseen sites appear. CipherSight learns representations from TLS records rather than raw TCP packet sequences, jointly encoding several record-level attributes in a hierarchical architecture that models both dependencies within a flow and interactions across concurrent flows, trained with a masked record modeling objective plus record-to-resource annotations used as privileged supervision. It reaches 95.41% accuracy over more than 2,000 website classes in the closed-world setting and stays above 90% accuracy under both temporal and geographic distribution drift, outperforming all evaluated baselines.

Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact

Zhelun (Allen), Wu Text-to-SQL systems built on language models fail in a particularly dangerous way for production use: a hallucinated column or mis-aggregated total returns a fluent wrong number that looks identical to a right one, especially when the consumer is a dashboard or a tool-using agent that never inspects the query. The proposed architecture pairs a generative shell, which interprets underspecified input and phrases replies, with a deterministic kernel that matches fully specified questions against a bounded set of answerable question shapes and compiles them to queries, meeting at a user-read confirmation before any value is computed. The governing invariant is that a component able to fabricate may influence which question gets answered but never which value gets returned, so unanswerable requests are declined because they are structurally unrepresentable rather than filtered by a confidence estimate; the pattern is worked across three domains and backed by a two-year production case study compared against a fine-tuned parser and a tool-retrieval agent.

Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy

Liwei Deng, Jing Jiang, Zhiwei Li, Yang Wang, Guodong Long Short-video recommenders optimized for engagement systematically favor content that grabs attention within seconds, and there is growing evidence that heavy exposure to such shallow content affects cognitive engagement and well-being. The proposed Content Depth Score rates how far a video is expected to stimulate higher-order cognitive processes on a seven-level scale grounded in cognitive psychology and learning theory, and SCOPE-Bench applies it as annotations over 150K videos from a large open-source dataset. Evaluating 13 representative recommenders reveals a consistent pull toward shallow content, and algorithms that do surface cognitively deep videos are only marginally better than random selection.

Buy the Rumor, Sell the News: When Is News Priced In?

Alireza Kargarzadeh, Nariman Khaledian, Navid Parvini, Sid Ghatak, Arman Khaledian Two market adages hold that news is already priced in by publication and that traders buy the rumor and sell the news, both placing the price move before rather than after publication. The test covers 4.57 million financial news articles on roughly 3,000 US stocks from 2023 to 2026, with a large language model teacher distilled by active learning into a compact classifier that assigns 17 event tags and five attributes, articles clustered into stories to distinguish first reports from follow-ups, and beta-adjusted abnormal returns measured around 1.68 million stock-day events against 364,405 neutral-sentiment placebo events. The cumulative move in the news direction by the close of publication day is 2.8 times its value 20 days later, and for rumor-flagged events the rumor day captures the entire move while confirmation adds nothing; quantified fundamentals such as earnings and guidance keep drifting for weeks while soft story-driven news gives its move back, and publicity raises volatility before publication and lowers it after.

HERMES: a multi-agent framework for structured knowledge extraction from ultra-long documents in geoscience

Ziqi Song, Zongyuan Xiang, James G. Ogg, Bruce S. Lieberman, Gabi Ogg, Natalia L\'opez Carranza et al. Much authoritative geoscience knowledge sits in decades-old monographs whose unstructured prose and complex layouts block computational reuse. HERMES is a multi-agent extraction framework in which a coordinating language model applies domain constraints, validation rules, and evidence tracing over parsed text, tables, figures, and captions of ultra-long documents. Applied to the 55-volume Treatise on Invertebrate Paleontology, it produced a public database of 32,277 fossil taxonomic entities and 451,878 attributes at roughly 0.90 F1 for entities and 0.91 for attributes, about six times faster per volume than the manual baseline, and transferred without retraining to palaeomagnetism and geochemistry documents.

Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency

Junchi Liu, Ali Bigdeli, Roya Daneshi, Atu Ambala, Sudipto Ghosh, Fabio Santos cross-listed Studies of language-model-based automated program repair mostly report aggregate fix rates, leaving open how performance depends on how complicated a bug is, how precisely the fault was localized, how much reasoning effort is spent, and what it all costs. Two repair techniques, ChatRepair and CodeCorrector, are run across DeepSeek, GPT, and Llama models under varying bug complexity and localization precision with statistical analysis of the results. Over half of moderately complex bugs are repaired by low-cost models, imprecise fault localization widens the gap between repair techniques substantially, and spending more — pricier models or heavier reasoning settings — does not reliably improve cost-efficiency: GPT-5 fixes 7 and 39 more complex bugs than DeepSeek-V4-pro and DeepSeek-V3.2 respectively, while DeepSeek-V3.2 is the most cost-efficient overall.

When Denoising Hurts: Rethinking the Terminal Step of Diffusion Time Series Forecasters -- Extended Version

Dat Nguyen-Cong, Luong Tran, Tung Kieu Diffusion forecasters are usually run to the end of the reverse process on the assumption that every denoising step refines the prediction. Tracking forecast quality across the reverse trajectory shows the opposite at the tail: broad temporal structure is recovered while noise is still high, and continued low-noise refinement introduces statistical drift that degrades accuracy — which also explains why prior work gravitated to narrow architecture and schedule choices. The proposed fix is a label-free global stopping criterion that detects when to terminate, plus a Bernoulli timestep sampler that concentrates training on the high-noise region where inference now ends while still covering the full schedule; experiments on eight real-world datasets show both faster inference and better accuracy.

Forecast Collapse in Time-Series Foundation Models

Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen et al. Forecasting hourly returns for 1,000 US equities makes time-series models emit nearly flat predictions that rank stocks poorly, a failure the authors name forecast collapse and that mostly vanishes when the target is trading volume instead. Sweeping time-series foundation models (TSFMs), twelve deep forecasters, and 97 benchmark configurations traces the effect to target predictability plus per-series training objectives that never identify cross-series structure, exposing a calibration-versus-ranking tradeoff: squared error flattens forecasts, while optimizing cross-sectional correlation directly can inflate amplitude by over an order of magnitude. Their CalibRank objective balances the two and nearly triples cross-sectional correlation on Finance1K while keeping amplitude near the target, improving correlation on every model tested. The broader point is that per-series evaluation metrics can hide the cross-series failures that downstream decisions depend on.

From Fixed Grids to Moving Particles:A Transferable Latent Operator for Fluid Dynamics

Meng Li, Chuqi Chen, Zhengqing Gao, Xi Zhou, Xiao Sun, Yang Xiang et al. Fluid simulation data is usually available as Eulerian fields on fixed grids, while Lagrangian particle trajectories — needed to describe transport — are scarce, so neural operators are almost always trained and tested in the Eulerian view. The Transferable Latent Operator (TLO) separates latent flow evolution from coordinate-dependent decoding: querying the evolving latent state at grid points yields Eulerian fields, while querying velocities at particle positions and integrating them forward produces Lagrangian rollouts from the same model. Across five fluid-dynamics benchmarks TLO beats existing neural operators on Eulerian prediction and generalizes zero-shot to Lagrangian particle rollout without any Lagrangian supervision, improving further with a small amount of Lagrangian fine-tuning.

BGA: A noise-immune neural distillation framework for malicious signature extraction in high-entropy encrypted flows

Sheng Hong, Yixuan Huang, Weiwei Jiang, Junyuan Zhang, Jiacheng Wang, Ruijian Jiao cross-listed Encrypted TLS 1.3 traffic is high-entropy enough that attention-based detectors spread their weight across cryptographic noise instead of attack signal. BGA first uses analysis of variance to separate discriminative control-plane features such as industrial setpoints from that noise, then trains a Wasserstein GAN with gradient penalty to synthesize rare-class flows within an 86,878-record corpus, and finally combines a bidirectional LSTM with an adaptive gated multi-head attention layer that acts as a filter on encryption artifacts. Recall on rare Malicious State Command Injection attacks rises by 43.2%, all key metrics exceed 95.2% on the CIC-IDS-2018 and Edge-IIoT benchmarks, and noise-injection tests give an 8.57% robustness margin over a vanilla Transformer at 0.282 ms inference latency, which the authors extrapolate to roughly 1.7 ms on ARM edge gateways.

MINT: A Universal Zero-Shot Predictor for Transaction Data

Parameswaran Kamalaruban, Viktor Drobnyi, Maeve Madigan, Julia Rozanova, David Sutton, Stuart Burrell Banks turn transaction sequences into embeddings with payments foundation models, but those embeddings feed task-specific heads and cannot answer novel questions zero-shot, whereas LLM-based approaches lose predictive signal and pay heavily for serializing transactions into text. The Multimodal Instruction Network for Transactions (MINT) wires a pretrained transaction sequence encoder into a decoder-only LLM through lightweight embedding injection, transaction-language alignment, and instruction tuning. It reports state-of-the-art predictive question-answering both in-distribution and out-of-distribution while substantially cutting input tokens, latency, and memory relative to text-serialization baselines, and analyses of representations, alignment strategies, training data, and history length support compact embeddings over serialized text for this kind of multimodal reasoning.

How Much Do Legal RAG Systems Still Hallucinate?

Souvick Das, Sallam Abualhaija, Domenico Bianculli Ungrounded answers carry real consequences in law, so this study audits hallucination in eight legal retrieval-augmented generation (RAG) systems over two corpora: the GDPR in English and a national civil law in French. Evaluation runs at both claim and answer level, reporting hallucination density and severity broken down by question category and user persona, with findings validated on 142 questions written by legal experts. Hallucination remains pervasive, affecting under 10% of responses for the best systems but close to half for the worst, and false-premise questions — those built on incorrect assumptions the system should reject — trigger especially high rates on the expert-authored set.

Non-Parametric Spatiotemporal Trajectory Prediction via State-Conditioned Transition Sampling

Michael Fore, Akshay Jain, Justin Downes, Rohan Pradhan, Duncan Botti Multi-modal trajectory prediction is normally handled by trained sequence models, which need substantial historical data and GPU training before they can serve a new geographic region. The alternative here is entirely training-free: a lookup table of historical state-to-next-position transitions, queried with a product kernel over spatial proximity, bearing, speed, and temporal context, with two inference modes over the same table — diversity-penalized sampling for covering distinct plausible routes, and beam search for the single most likely path. On the TrAISformer benchmark of Danish maritime AIS data it matches a 57M-parameter transformer at full data availability with zero learned parameters and no GPU, and stays stable down to 10% of the training data where the transformer degrades catastrophically, which points to deployment in new regions from an order of magnitude less history.

Catching the Imposter: Self-Supervised Learning of Physical Coherence with Cross-Entity Feature Permutations

Aleksei Rozanov, Arvind Renganathan, Vipin Kumar Scientific measurements describe entities whose variables are coupled by physical law, a structure that standard self-supervised objectives leave unused. The imposter pretext task swaps a subset of one entity's feature values with genuine observations donated by a different entity and trains the encoder to spot which features were swapped; since every donated value is individually plausible, the only way to solve it is to learn cross-feature physical dependencies. Evaluated on 21 variables from global ERA5-Land reanalysis across seven downstream tasks in climate classification, carbon flux estimation, and streamflow prediction — the first systematic comparison of self-supervised objectives for land-surface modelling under a shared architecture and pre-training budget — the results show no single objective dominates, with the best pretext task depending on the downstream family and imposter contributing information complementary to existing objectives.

Universal Thermodynamic Interatomic Potentials for Crystalline Materials

Juno Nam, Bowen Deng, Xiaochen Du, Luis Barroso-Luque, Benjamin Kurt Miller, Rafael G\'omez-Bombarelli cross-listed Solid-state phase stability is governed by free energies, but high-throughput materials screening still leans on ground-state energies because free energies require ensemble averages that are expensive to compute. The thermodynamic interatomic potential (TIP) extends a standard interatomic potential from static energy to a thermodynamically consistent Gibbs free energy model, so responses to temperature and pressure follow by automatic differentiation; the implementation TIP[UMA] builds on the universal potential UMA, trains on free energies spanning quasi-harmonic to molecular-dynamics fidelity, and is calibrated against higher-resolution calculations or experiment. A single evaluation returns a crystal's equation of state and locates phase transitions among competing branches, including dynamically stabilized phases, and fine-tuning extends coverage to alloy solubility limits and miscibility gaps.
51 more specialized papers

Large Language Models 45

Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking

Pradeep Kumar Sharma, Shantanu Godbole, Hritvik Shrivastava Which layers of a sparse Mixture-of-Experts (MoE) model can be pruned is poorly characterized, so the authors mask low-magnitude experts layer by layer in Qwen3.6-35B-A3B (40 MoE layers, 256 experts each, top-8 routing) and score output quality on the XLCoST cross-lingual code-translation benchmark at 100-, 300- and 500-prompt scales. Sensitivity proves strongly depth-dependent: early and middle layers degrade badly under masking, while the last five layers tolerate aggressive removal of low-magnitude experts. Masking 640 of 10,240 experts confined to layers 35-39 retains 419/500 good-or-similar outputs, against 150/300 for flat 30% masking across all layers. Narrowing routing from top-8 to top-6 active experts cut wall-clock time with no quality loss on a small probe, but did not compose cleanly with heavy masking.

Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required

Egor Shibaev, Vera Kudrevskaia, Timur Galimzyanov, Mikhail Evtikhiev, Ana Terna, Rastislav Rabatin et al. Post-training reports, model cards and blog posts routinely treat SWE-bench and LiveCodeBench scores as evidence of general coding ability, and the authors build a Django-based multi-task benchmark suite to test whether that inference survives contact with different tasks. Evaluating foundation models alongside checkpoints post-trained on SWE-bench trajectories, they find benchmark rankings frequently fail to generalize, with SWE-bench optimization yielding limited or no gains on the Django tasks or on LiveCodeBench, and fine-tuning on individual Django task types likewise failing to transfer. Their recommendation is differentiated evaluation — holistic assessment for frontier models, multi-task suites for research, human-in-the-loop studies for narrow applications — plus a capability taxonomy and sustained benchmark maintenance instead of one-off releases.

Modular Cognitive Architecture Emerges in Large Language Models

Pengrui Han, Jacob Andreas, Evelina Fedorenko, Andrea Gregor de Varda Human brains show pronounced functional specialization, with separate networks handling language, formal reasoning, reasoning about other minds, and reasoning about the physical world; the question here is whether that modularity is a necessary property of intelligent systems or an accident of biology. Circuit analyses across 46 tasks spanning those four cognitive domains trace which neurons each task recruits inside large language models. Tasks that share a brain network in humans recruit overlapping neurons in the models, while tasks drawing on different human networks recruit distinct ones, which the authors read as modular organization emerging convergently under a very different optimization process.

A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing

William Nixon, Jon Durbin, Florian Standhartinger, Haryadi S. Gunawi, Juncheng Yang Studies of production language-model serving so far cover short windows and few models, leaving little visibility into how traffic evolves or how user-model interaction shapes it. The authors characterize a full year of production traffic from Chutes, spanning many models and users including long-tail ones, analyzed from aggregate, temporal, model-level and user-level perspectives with attention to caching and load-balancing behavior. The complete one-year trace will be released alongside the paper so serving-systems research can benchmark against real rather than sampled or synthetic workloads.

Jais 2: A Family of Arabic-Centric Open Large Language Models

Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal et al. Jais 2 is a family of Arabic-centric open models from MBZUAI, Cerebras and Inception, comprising what the authors report as the largest open Arabic-centric model trained from scratch at 70B parameters plus a competitive 8B variant. A custom Arabic-centric vocabulary together with an optimized architecture and training recipe let the models reach strong Arabic results on a substantially smaller token budget than comparable systems, leading the evaluated open models on OALL2 and AraGen while staying competitive in English and performing well on culturally grounded benchmarks covering poetry, religion, cuisine and dream interpretation. Weights are released on HuggingFace under a commercially permissive license, and the 70B chat deployment runs on Cerebras hardware at up to 2,000 tokens per second.

IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering

JungMin Yun, YoungBin Kim Multi-hop question answering buries retrieval-augmented systems in long, noisy contexts, and existing prompt compressors are designed for single-turn queries so they discard evidence that only becomes relevant at a later reasoning step. IterCOMP is training-free: it decomposes retrieved documents into evidence segments, judges whether the question is answerable yet, and generates targeted follow-up questions to pull in missing evidence over successive rounds, ending with a compact reasoning-oriented prompt. On MuSiQue, 2WikiMultiHopQA and HotpotQA it improves exact-match and F1 over existing compression baselines while cutting the token budget, with the advantage growing as reasoning complexity increases.

Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors

Akira Okutomi Confidently wrong answers from language models are usually read as a sign that the underlying inference is shaky, but this work tests whether some are instead locally stable — unchanged by small perturbations. Two diagnostics are combined: an output-level audit score that ranks domains by confidence variation and overconfident mistakes under a forced-answer baseline, and an internal probe measuring how far hidden states move. Self-critical prompting reliably reduced hidden-state sensitivity across layers in three open-weight models, yet overconfident errors were not measurably more locally sensitive than confidently correct answers, implying that prompting stabilizes representations without actually fixing calibration.

Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning

Jiahe Fan, Si Chen, Yinghao Hou, Aiyuan Zhang, Hong Xie Transferring capability from a large donor model into a much smaller recipient normally requires aligning neurons or architectures, which is hard when the two differ in scale and structure. Activation-Prune-Merge (APM) instead builds task-conditioned activation maps on the donor, uses them to select the most salient layers, hidden dimensions, attention heads, and MLP neurons, prunes the donor down to the recipient's shape, and blends the resulting slice in with a very small interpolation coefficient — no training required. Across 16 benchmarks covering reasoning, math, code generation, instruction following, and classification, average accuracy of a 3B recipient rose from 55.5% to 60.6%, with RTE climbing from 64.3% to 82.3% and BoolQ from 70.8% to 79.2%.

No Universal Signal Predicts Sample-Level LLM Regression under Version Updates

Jia Sheng, Yiwei Lu A new model version usually beats the old one on aggregate scores while still breaking individual cases that previously worked, and the question here is whether such per-sample regressions can be predicted at inference time. Single-model signals (confidence, logit margin, attention entropy) are pitted against cross-version signals (output and token-level KL divergence, likelihood drift, representation drift) under a common test that isolates each signal's gain over a plain confidence baseline, spanning six benchmarks in multiple-choice question answering, math reasoning, and code generation across six model update pairs. No signal won universally — confidence dominated on multiple-choice and easy math, while likelihood and KL divergence signals helped most on harder math and code — but some cross-version signals stayed informative without labels where confidence failed, enabling a proof-of-concept fallback that routes high-risk inputs back to the previous model.

The Query Knows What to Forget: A Second Erase Direction for Linear Attention

Dhruman Gupta, Aritra Das, Debayan Gupta Linear attention compresses history into a fixed-size state, so at long context many stored items crowd each other and retrieval degrades. Delta-rule models including Gated DeltaNet-2 derive their erase vector from the current token's key, but read interference is measured through the query, a direction the key-based erase cannot touch. The Query-derived Erase Direction (QED) adds a second erase component taken from the query and made orthogonal to the key, using the part of the state a key-directed edit cannot reach to cancel stale content along the query; it improves retrieval at every length beyond the training window and roughly doubles usable context length on S-NIAH-1.

From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models

Pranav Kumar Kaliaperumal A retrospective covering October 2018 to July 2026 tracks how language models went from BERT to agents that solve competition math and write software, reporting that the ability to resolve real coding issues improved nearly sixfold per year since late 2024 while cost per unit of capability collapsed, with budget-tier models matching flagship capability at one to six dollars per million tokens. It also documents fragmentation of the frontier into task-targeted models, with different systems leading on frontend coding, repository-level coding, and terminal tasks. A companion grade-school math experiment on Qwen2.5 shows basic decoding solving 58 of 100 problems against up to 79 with advanced sampling, and a confidence ranker placing 47 correct answers in its top 50, with all materials released publicly.

Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT

Pu Zhao, Changdi Yang, Yixiao Chen, Yi Gao, Yifan Cao, Haochen Zeng et al. cross-listed Translating legacy C into safe, idiomatic Rust would remove whole classes of memory-safety bugs, but off-the-shelf language models are weak at it because general pretraining emphasizes neither idiomatic Rust nor cross-language semantic equivalence nor repairing compiler feedback. The reported recipe applies a three-stage curriculum to Qwen3-27B: continued pretraining on Rust-heavy corpora, supervised fine-tuning on microsoft/Verus_Training_Data to instill debugging and self-repair, and task-specific fine-tuning on paired C/Rust LeetCode solutions. Evaluation runs through SACTOR, an agentic static-analysis-guided framework that does two-phase unidiomatic-to-idiomatic translation with foreign-function-interface end-to-end tests, reporting success rate, Clippy lint counts and unsafe-code fraction as idiomaticity measures, and failure-mode breakdowns against baseline Qwen3-27B and other models.

Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline

Jo\`ao Pedro Monteiro Pereira, Vinicius Cardoso Garcia cross-listed Non-functional requirements in code-generation prompts are usually a terse one-line phrase, and it is unclear whether spelling them out against a quality standard helps. Four requirements — performance, error handling, code smell, readability — were expressed three ways (a one-line baseline, rich prose grounded in the ISO/IEC 25010 quality model, and structured JSON with the same ISO content) across ten prompt variations each, evaluated on HumanEval and HumanEval-ET under a fixed model snapshot with paired non-parametric tests. ISO grounding improved static quality proxies and reduced sensitivity to prompt wording but did not reliably improve functional correctness, and for error handling the extended-test pass rate actually fell, suggesting defensive coding conflicts with exact-output benchmarks. Holding ISO content constant, prose and JSON differed negligibly in correctness, so the semantic content matters more than the serialization format.

The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference

Teng-Ruei Chen Two GPU kernels implementing the same scaled INT8 general matrix multiply interface are normally assumed interchangeable, so swapping only the INT8 linear kernel (CUTLASS versus Triton) inside vLLM — with checkpoint, prompts, hardware, decoding and quantization config held fixed — should not change output. Each arm reproduced itself bit-for-bit across restarts, yet the two arms agreed on no generated sequence in any end-to-end comparison run (0 of 8, 0 of 16, 0 of 64), even though the INT32 accumulator is provably exact and order-independent under a verified no-overflow bound. Feeding both kernels identical operands from every linear layer of Qwen3-1.7B and 8B produced bit-identical outputs under power-of-two scales and differences of at most one bfloat16 spacing under the real scales, localizing the divergence to scale application and output rounding after the accumulator; applying that as a probe restored full bitwise agreement end to end. Cross-implementation FP8 shows a different signature that grows with reduction depth, and teacher-forced replay shows flips concentrate at small logit margins, predicting flip risk with an ROC-AUC of 0.94.

Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future Directions

Qinglin Yang, Chen Qiu, Hongyuan Zhang, Pengdeng Li, Yuan Liu, Zhihong Tian Federated learning lets clients train together without pooling raw data, and combining it with prompt-based adaptation offers a way to specialize large language models without the compute or data centralization that full federated fine-tuning requires. This survey organizes federated prompt learning across the model lifecycle — pre-training, fine-tuning, and deployment — covering how it differs from conventional federated learning and full-model federated fine-tuning, and cataloging defense mechanisms for the attacks it inherits. The comparison centers on the trade-offs each approach makes among accuracy, communication cost, computational overhead, scalability, personalization, and client heterogeneity, closing with open security, privacy, robustness, and systems challenges.

Bootstrapping Niche Multilingual Code Translation via Reinforcement Learning with Execution-Based Verifiable Supervision

Kouki Yuki, Jie Zeng, Kyoko Ogawa, Ryunosuke Ikeda, Yohei Kobashi, Takeshi Kojima et al. Neural code translation has concentrated on a few popular languages such as C++, Java, and Python, leaving many-to-many translation among less common languages short of parallel supervision and prone to producing plausible but non-executable code. The pipeline grows verified seed Python programs into an execution-validated multilingual pool, labels candidate translations by whether they actually run, trains a reward model on the resulting preferences, and optimizes base models with GRPO over 600 directed language pairs. A new execution-based benchmark, HumanEval-X++, extends HumanEval-X to this wider language space, and the 4B Qwen-3.5 model improves by 13% on average across all languages, with a 21% gain on mid-tier languages.

Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification

Benjam\'in Schindler, Gonzalo A. Ruz Synthetic training text from large language models varies widely in usefulness: some generated samples land inside the correct class region of an embedding space while others drift to the periphery or into other classes. The proposed filter scores each generated sample by its Euclidean distance to real class examples in a sentence embedding space and converts those scores into soft sample weights for classifier training. Across 13 datasets, 5 classifiers, 10 augmentation methods, and more than 6,700 configurations, the method beats SMOTE by 2.61 percentage points with an 88.9% win rate and transfers unchanged to named entity recognition for a 9.26-point gain; notably, the simplest distance-based filter consistently outperforms more elaborate multi-criteria variants.

CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

Yuji Ren, Chenkai Xu, Zhuocheng Gong, Jianguo Li, Zhijie Deng Diffusion language models generate faster by filling several masked positions per forward pass, but aggressive parallelism makes their early denoising predictions unreliable and those errors compound downstream. Consistency Forcing is a distillation method that trains on pre-collected self-rollout trajectories so early-stage mask predictions align with later-stage ones, using a confidence-adaptive Kullback-Leibler objective that blends the strengths of forward and reverse divergence, with theory explaining why the consistency target approximately minimizes early-stage prediction error. The same formulation covers both mask-to-token and edit-capable decoding, and experiments on LLaDA variants show improved speed-quality trade-offs that widen under high-parallelism decoding budgets.

Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion

Hwan Chang, Yongil Kim, Heuiyeen Yeen, Yireun Kim, Jinsik Lee, Hwanhee Lee Creative-writing training data for large language models is dominated by stories, so models trained on it struggle with the structural and formatting conventions of other creative formats. The proposed framework separates thematic breadth from genre control by pairing human-authored story prompts as creative seeds with manually curated genre attributes that enforce distinct structure and style, then prompts strong models for query-response pairs and quality-filters them, producing the Multi-Genre Collection of 50K examples across 13 genres including rap, lyrics, scripts, game design, and character design. Models fine-tuned on this corpus beat base models, writing-specialized baselines, and models trained on existing writing corpora on out-of-distribution benchmarks, and genre-count ablations indicate that controlled genre expansion, not more story data, is the main driver of the gains.

Repair, Not Improvement: Decomposing Constrained Decoding in Tool-Call Abstention

Janghoon Lee (Redrob) Prior accounting of constrained generation found the decoder contributes little for format constraints but explicitly warned against extrapolating to cases where a constraint encodes correctness, naming function calling as one; this study measures that excluded case, focusing on when a model should decline to call any tool. Three conditions over one byte-identical prompt separate a grammar's two jobs — fixing where generation stops and which tokens may be emitted — across open-weight models from 0.6B to 4B on matched English and Korean items. Against an unconstrained decoder, constrained decoding is negative on abstention in four of six cells with intervals excluding zero, worst -29.5 points, and positive in none, and the total is a sum of opposing effects: on the smallest model in Korean the stop token costs -20.0 while the enum returns +19.5. What the grammar recovers is form rather than judgment, since 545 of 698 repaired abstentions had no readable answer to begin with, and both preregistered language claims fail.

QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

Vincent Counathe, Ben Athiwaratkun, Christopher De Sa, Tianyi Zhang Quantization-aware training (QAT) computes losses and surrogate gradients from a lossy reconstruction of latent full-precision weights while updating the latent weights themselves, a mismatch that raises the final loss floor; second-order post-training methods close a similar gap but take hours per pass and cannot be repeated as weights evolve. QUASAR performs lightweight loss-aware reconstruction inside the training loop, using an exponential moving average of squared gradients as online saliency estimates, searching a small set of clipping ranges, and fitting affine dequantizers by saliency-weighted least squares; the accompanying analysis shows loss-aware reconstruction error is the only reconstruction-dependent term in the QAT convergence bound. It changes only training, keeps standard deployment formats including integer quantization and NVFP4 with no inference overhead, and on Qwen3 and Llama-3.1 cuts held-out KL divergence by at least 10% at 3 and 4 bits and 29% at 2 bits, improving average accuracy across eight tasks by 3.5 to 4.3 percentage points at 2 bits.

Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead

John T. Halloran Nanbeige4.2-3B is a 3B-parameter agentic model whose Looped Transformer reuses one layer stack for a second forward pass, buying effective depth without extra parameters but roughly doubling peak attention memory. Running it on Apple Silicon surfaced five independent bugs blocking the released checkpoint under Hugging Face transformers, including a silently zeroed rotary position embedding buffer and calls to removed cache APIs; a chunked-prefill strategy then addresses the memory penalty, extending allowable context width by 2.7x on 32 GiB of shared memory. After further system-prompt and Metal-backend memory patches, the debugged model completes up to 30% of real agentic tasks on a subset of MCPMark (up from 0%) and is near-perfect on single tool calls in BFCL while failing most multi-tool tests; the patched checkpoint and harnesses are released.

Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model

Yongmin Kim, Shota Takashiro, Yusuke Iwasawa, Takeshi Kojima, Yutaka Matsuo Large reasoning models generate long chains of thought that are expensive to serve, and production throughput requires batching, but existing training-free adaptive pruning methods collapse in that regime: a batch shares one pruning mask, so activations are aggregated across samples while the pruning threshold was calibrated offline on unaggregated activations, causing the realized sparsity to drift. The proposed method replaces threshold selection with periodic top-k selection over aggregated importance scores, which is immune to the distribution shift aggregation induces and runs once per update period rather than per token, and adds an activation memory that accumulates importance across phases because important neurons re-fire periodically during long generations. On DeepSeek-R1-Distill-Qwen-7B at batch size 4 and 50% target sparsity it beats the previous state-of-the-art adaptive pruning method by 39.7 percentage points in average accuracy, reaching a 1.40x speedup over dense inference at 50% actual sparsity.

Scaling Domain Data Repetition in LLM Pretraining

Jingwei Li, Xinran Gu, Rui Dai, Xintong Hao, Chengyin Xu, Yan Wu et al. As models grow, token budgets grow with them to keep a sensible tokens-per-parameter ratio, but scarce high-quality domain data cannot scale the same way and gets diluted in the mixture; repeating it counteracts dilution at the risk of overfitting. Sweeping this trade-off under practical scaling, where token budget grows with model size, yields two findings: at fixed tokens-per-parameter the optimal repetition count mildly increases with model size, and across domains the optimal count correlates strongly and negatively with a domain's final validation loss, while the amount of unique domain data barely matters. The practical implication is that repetition counts tuned on small proxy models at the same tokens-per-parameter ratio carry over to larger runs.

The conditional superiority of fast silicon sampling

Nickolas Hock Yuen Lam, Ji Xuan Voo, Xiangyu Ma Silicon sampling — using language models to stand in for human survey respondents — sometimes reproduces population statistics well, and the question here is whether a cheaper, faster sampling mode sacrifices that fidelity. Fast and slow modes are compared against a nationally representative sample of Singaporean respondents. Both modes estimate population means moderately well but understate opinion variance and distort the latent contextual structure behind human views, so the method warrants caution overall; conditional on those limits, the fast mode is uniformly at least as faithful as the slow mode while using far less compute and wall-clock time.

P2Skill: Privacy Preserving Skill Distillation for Cloud-Local LLM Inference Systems

Myunghoon Ryu, Geunpyo Park, Sungjoon Lee, XinYu Piao, Jong-Kook Kim cross-listed Splitting inference between a phone-resident small model and a cloud model only protects users if cloud-bound requests are stripped of personally identifiable information (PII), and existing masking or perturbation schemes distort meaning or demand extra training. P2Skill instead hands a local small language model (SLM) a set of prompt-defined skills — query decomposition, PII-aware routing, paraphrasing, and reconstruction — that a cloud LLM iteratively rewrites based on observed execution failures, so the system needs no privacy-specific fine-tuning and no learned PII detector. On a four-domain benchmark it reports 1.69× and 3.66× higher privacy-preserved inference quality than prior baselines.

Retrieval Grounding Latent Reasoning for Dense Retrieval

Gang Zhou, Xiongxi Yu, Hu Tian, Yang Wei, Lu Pan, Ke Zeng et al. Reasoning-intensive retrieval needs embeddings that encode not just topical similarity but the inference required to judge relevance under an instruction, and models that bolt reasoning onto embeddings are typically trained end-to-end on the retrieval loss alone — leaving room for shortcut latent trajectories that score well without adding anything. RGLT (Retrieval Grounding Latent Reasoning) runs non-autoregressive reasoning in hidden space over instruction-conditioned silent tokens, shaping intermediate states with stage-wise chain-of-thought reconstruction distilled from explicit reasoning and assigning retrieval-effect credit so each latent step is optimized for its incremental retrieval gain. On reasoning-intensive retrieval benchmarks it beats strong baselines while keeping embedding inference cheap.

KV Cache Compression Through the Lens of Transform Coding

Hannah Laus, Claudio Mayrink Verdun, Hao Wang, Flavio du Pin Calmon, Felix Krahmer The key-value (KV) cache dominates memory in long-context inference, and existing quantization schemes minimize reconstruction error in the cache itself without asking how that error propagates through attention. Under a white-noise quantization model the authors prove the expected attention-aware distortion splits into additive key and value terms that factor across tokens and channels, which lets them apply transform coding and reverse water-filling from rate-distortion theory to allocate bits over a calibration set. The resulting Attention-Aware Transform Coding (AATC) reaches near-lossless accuracy at about 5.8× compression on Llama-3.1-8B-Instruct and Qwen-2.5-7B-Instruct across LongBench, RULER, GSM8K, MMLU-Pro, and MATH-500, while every baseline degrades in at least some settings.

FreeBalance: Pre-Routing Online Moe Load Balancing via Residual Workload Prediction

Pengfei Chen, Yize Wu, Shouxu Kuang, Ke Gao, Ling Li In distributed Mixture-of-Experts (MoE) inference, skewed routing leaves one rank overloaded and stalling everyone else, and existing online rebalancers can only act after each router has produced its decisions, putting expert weight migration squarely on the critical path. FreeBalance predicts the routing distribution in advance from cross-layer similarity of hidden representations in the residual stream, so expert migration can be planned early and overlapped with pre-routing computation such as attention, with a cost model capping the number of swaps so synchronization stays hidden inside the available window. Across models and datasets it cuts the max-to-mean rank load ratio by 32.8% and end-to-end prefill latency by 13.1%, hiding migration of an average 5.1 experts per layer that would otherwise cost about 8.5% of critical-path latency.

APTER: Adaptive Post-Training with Expert-Grounded Rubrics

Xukai Wang, Liangqi Li, Zhiyue Xu, Jingang Zhou, Xiaoyu Shi, Jiansheng Cai et al. Professional-domain deployments need models that respect domain constraints, cite critical evidence, and reason completely, but holistic preference tuning and outcome-level verification give no purchase on those requirements, and per-query generated rubrics tend to drift and skip essentials. APTER grounds rubrics in an expert criteria framework where each criterion names a stable professional capability, selecting and instantiating relevant criteria per query to produce executable supervision without reference answers. Rubric verdicts then serve double duty during reinforcement learning: as the optimization signal, and — when low scores are aggregated by criterion identifier — as a diagnosis of persistent weaknesses that triggers targeted supervised fine-tuning updates. Across three model generations, averages improve over the base models by up to 15.86 points on mathematical reasoning and 8.04 on medical question answering.

Multi-Objective Bayesian Optimization for Model Merging

Utkarsh Agarwal, Vamshi Bonagiri, Raul Astudillo, Monojit Choudhury Merging trained models in weight space avoids extra fine-tuning, but choosing merge coefficients is hard because evaluations are expensive, there are no gradients, and source capabilities trade off against each other. MOBO-Merge casts coefficient selection as black-box multi-objective optimization and uses Bayesian optimization to approximate the Pareto front within a fixed evaluation budget, independent of which merge operator is used. Testing Qwen3-4B and Llama-3.1-8B across two-model instruction-math and three-model instruction-math-code settings with Linear, SLERP, TIES, and block-wise operators, it beat random search on held-out hypervolume in 11 of 12 comparisons, with negligible benefit for one-dimensional linear interpolation and substantial gains for the higher-dimensional TIES, block-wise, and three-objective searches. No single operator dominated: TIES led three of four family-setting combinations while Block-Linear 4x won the Llama three-model merge.

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

Haonan He, Haodi Lei, Yun Luo, Haoran Zhang, Shunkai Zhang, Yizhuo Li et al. On-policy distillation from a long-context reasoning teacher into a short-context student breaks down over tokenizer mismatch, distribution mismatch, runaway response length, and unstable training. SimpleOPD performs distillation in a shared text space, aligning only tokens that cover identical text spans under both tokenizers, and controls length by adding a student reference KL term plus masking the advantages of termination tokens such as </think> so the student grows its reasoning steadily instead of drifting and getting truncated. Transferring proof-reasoning ability from SU-01 to Qwen3, Qwen3.5, Intern-S2, GLM-4.7, and Gemma-4 students improved mathematical reasoning across both same-family and cross-family pairs, with Intern-S2-Preview gaining 21.2 points on ProofBench to reach 55.2 and surpassing Gemini-2.5-Pro, alongside gains on science benchmarks HLE and HiPhO that suggest transfer beyond the training domain.

Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

Kai Chen, Jifeng Ding, Ning Ding, Jiaye Ge, Lixin Gu, Yicheng Gu et al. Mobius-v0 splits a transformer's usual entanglement of knowledge and computation into a globally shared Memory built from feed-forward layers that stores knowledge vectors, and multiple Reasoners built from self-attention that repeatedly query that memory using hidden states as both cache and carrier. Because knowledge lives in one shared store while reasoning is applied iteratively, the architecture compresses knowledge more tightly and reuses reasoning operators. A 7B model trained from scratch matched a 7B transformer baseline's downstream scores using 62.6% of the baseline's training data, and Intern-S2-Mobius, continually pretrained from Qwen3.5-35B, matched downstream quality while delivering nearly 4x end-to-end inference speedup.

AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

Yiderigun Borjigin, Alexander Hermann, Christian Cyron, Roland Aydin Anchoring, the human bias in which an arbitrary reference number pulls a later numeric judgment toward itself, has been observed in language models, but prior studies test only a few ways of introducing the anchor and rarely separate obviously irrelevant anchors from plausible ones. AnchorBench evaluates multiple anchor pathways — including anchors injected through external context and retrieval-augmented generation — along an explicit relevance axis, across ten open-weight and four frontier API models. Susceptibility turns out to be strongly pathway-dependent, plausible anchors shift answers more than irrelevant ones when delivered through the stronger pathways, and influence fades as the anchor moves further from the evidence-supported answer. Notably, frontier models scoring above 95% on the anchor-free control condition remain vulnerable to plausible anchors, so task accuracy is no proxy for robustness.

Local and Global Regimes of Geometric Complexity in Language Model Representations

Arwa Osman, Marco Baroni, Iuri Macocco Intrinsic dimensionality is a common probe for how complex a language model's representations are, but it is unclear whether measured differences reflect language or artefacts of dataset construction. Holding everything else fixed and varying lexical diversity — the number of unique final tokens in a dataset — reveals a scale-dependent reversal: at low lexical diversity fewer unique final words give higher intrinsic dimension, while at high diversity the ordering flips. The authors derive an exact, parameter-free formula for the crossover point that matches the empirical transition at every scale tested, which both cautions against reading intrinsic dimension as a direct measure of complexity and describes an organising principle of the representation manifold.

DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding

Zewen Jin, Shen Fu, Zeping Duan, Shannon Wang, Weihao Wu, Chengjie Tang et al. Mixture-of-Experts (MoE) models serving latency-sensitive workloads such as coding assistants run at small batch sizes, where inference becomes memory-bound and the dominant cost is streaming expert weights rather than computing with them. DeaMoE restructures the expert layer into "departments" whose member experts share most of their parameters and keep only a small private residual, paired with a two-stage router that avoids loading redundant weights. Per-step loaded weights fall by up to 50.9%, giving up to 1.33x end-to-end time-per-output-token (TPOT) speedup for a pretrained 7B model on an A40, and microbenchmarks on DeepSeek-V3 show peak speedups of 2.00x on A40 and 1.97x on H100.

More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It

Haohui Yang, Jiaxing Sun, Xiujun Ma Power Sampling sharpens a language model's distribution over whole generation trajectories, a verifier-free way to push probability mass toward correct reasoning paths that is often proposed as a front end for other inference-time methods. Testing it with self-consistency reveals a paradox: it can increase the mass on correct trajectories while cutting downstream accuracy by as much as 18.5 percentage points across models and reasoning benchmarks. The authors attribute this to a dose mismatch, where a single fixed exponent changes different problems by wildly different amounts, and a coverage mismatch, where global sharpening collapses onto a few dominant paths — so a high pass@k can coexist with the loss of the broad path support that aggregation and search need. Replacing uniform trajectory exponentiation with a deformation-controlled, support-preserving target that calibrates sharpening per problem reverses the losses and, at equal budget with weighted self-consistency, beats standard multi-sample inference.

Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations

Toby D. Pilditch Benchmark runs typically sample every item the same fixed number of times, spending compute on items whose scores stopped moving many epochs ago. optstop reframes evaluation as sequential measurement built on hierarchical Bayesian inference: sampling continues where uncertainty is still high and halts where estimates are precise or stable, with support for binary, ordinal, and continuous outcomes, no need for a calibrated item bank, live or retrospective operation, and a safeguard that samples more cautiously near zero performance where rare successes carry the most information. On an illustrative 200-item, 10-epoch evaluation it eliminates 57% to 97% of planned trials across nine validation settings while reaching the same overall conclusions, with savings depending on how the evaluation is designed.

Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation

Isabel Cachola, William Walden, Reno Kriz, Mark Dredze Summarization metrics measure overall quality or specific properties like factuality, but not whether a summary actually serves a particular reader — a biomedical researcher and a family physician want different things from the same vaccine paper, and a short query rarely captures that gap. The argument is that a reader's persona, meaning their role and expertise, is more stable than any single query and recovers the missing context, so metrics should be tested for sensitivity to both informational and persona differences. Perturbation tests show that popular metrics including strong LLM-as-judge scorers fail basic checks on informational content, and an expert human study of persona-conditioned preferences finds that traditional and LLM-based metrics alike agree poorly with human judgments of information satisfaction.

You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model

Ziyang Luo, Zhongyao Chu, Xinjie He, Youting Wang, Xukui Qin, Runxiong Wu et al. A frozen language model both under-uses evidence already present in its residual stream and fails to notice when the input cannot support an answer, so it confabulates; two known fixes — a conditional steering probe that writes to mid-stack layers, and a zero-shot sufficiency direction that reads the stream to trigger abstention — interfere when combined, since steering shifts the very state the direction inspects. YOPO keeps the direction fixed and trains a small network to reconstruct the pre-steering residual from the steered one using mean-squared error on paired activations and no sufficiency labels, then reads the direction on that reconstruction, so answering, steering, and abstaining all happen in a single forward pass. On frozen Qwen2.5 backbones at 1.5B, 3B, and 7B, three-way accuracy more than doubles the frozen baseline (0.375 to 0.798 on alphaNLI at 1.5B) and the single pass beats the two-pass reference at every scale and across ten backbones from six families. A source-side audit caught a surface artifact leaking in the authors' own alphaNLI construction, so the architectural claims rest on native-label replications with SQuAD2, RepLiQA, and MuSiQue.

Approximate Muon with low-rank adapters

Ben Anson, Conor Houghton, Edward Milsom The Muon optimizer helps during pretraining but is seldom used for parameter-efficient fine-tuning, largely because LoRA's low-rank parameterization makes it mathematically impossible to orthogonalize the resulting weight update. The fix approximates a relaxed Muon objective in the low-rank setting via linearization followed by least squares, with an implementation that needs only matrix multiplications instead of costlier decomposition routines. Across supervised fine-tuning and a ReLoRA pretraining run, sMuon performs favorably, though the authors are explicit that results vary by model and evaluation and that the overall gain from Muon-style low-rank fine-tuning is moderate rather than dramatic.

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

Zhelun Wu Systems that ask a language model to draw a conclusion from many sources typically concatenate everything into one prompt, conflating two jobs with different needs: interpreting a single source rewards model capacity and context, while combining interpretations rewards fixed arithmetic, cross-instance comparability, and the option to return nothing. Splitting them turns the design problem into the interface, here a four-field evidence tuple of hypothesis, reliability bucket, rationale, and provenance — and exposes a failure mode the authors name count-scale drift, where thresholding a sum of unnormalized weights is posterior thresholding at an operating point that slides with how many sources were consulted, so no single threshold reconciles vote order with posterior order when source reliabilities differ. Pooling calibrated log-likelihood ratios fixes both problems arithmetically rather than architecturally, and applies equally to score-summing triage engines, diagnostic panels, and additive detectors. Instantiated twice on one longitudinal corpus, the separation helps in both settings, with a small sequence encoder on an auxiliary objective plus a tree ensemble carrying a censored survival loss reaching 0.921 AUPRC against 0.805 for a hand-crafted baseline; the paper also states five falsifying predictions, three negative results, and which comparisons remain confounded.

Handover of In-Context Learning State Across Session Boundaries

Masahiro Kato, Taka Kato When a long-running task outgrows a model's context window, the application restarts, or another agent takes over, something has to decide what information carries into the next session. The work formalizes this as transferring a task-relative in-context learning (ICL) state, separating exact recovery of the earlier text from preservation of the target predictive distribution, and shows that under an exogeneity condition predictive equivalence characterizes the coarsest sufficient handover and yields a fixed-length bit requirement. It proposes a three-part record — decisions and constraints stored verbatim, task-justified summary statistics for repeated evidence, and raw observations whose effect those statistics fail to capture — and quantifies the penalty for writing the record before the downstream query is known. Gaussian linear regression admits an exact finite-dimensional handover with finite-bit perturbation bounds, while nonparametric regression yields matching upper and lower bounds linking memory size to squared prediction error.
2 more specialized papers

Agents 39

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan Scaling agent evaluation usually means swapping an executable environment reward for a second language model as judge, and such judges — whether hand-written like G-Eval or fine-tuned — tend to credit fluent but unsuccessful trajectories as successes. RubricForge instead evolves the text of a judging rubric against a small set of ground-truth-labeled trajectories to maximize agreement with the real environment reward, then freezes it and applies it in one model call with no environment access, leaving a human-readable artifact whose verdicts trace to named criteria. Using a single 7B model as both agent and judge on tau-bench and WebShop, aggregate agreement with G-Eval is statistically indistinguishable, but the false-pass rate on failed trajectories drops roughly by half, 0.115 versus 0.173 on tau-bench — the quantity the authors argue matters, since a false pass ships a broken agent while a false fail only costs a retry.

Does a Language Server Save Tokens for Coding Agents? A Measurement Methodology and Preliminary Study

Pengcheng Xu Coding agents spend most of their context budget on retrieval, and the claim that typed semantic retrieval through the Language Server Protocol (LSP) is more token-efficient than lexical grep is, the authors note, asserted almost everywhere and measured almost nowhere. They formalize the question with a tokens-to-success metric, specify a five-arm ablation isolating semantic retrieval from confounds, and run a preliminary study on Python and TypeScript repositories with Claude Opus 4.8, Sonnet 4.6 and Haiku 4.5. The answer is conditional and usually negative: on symbol-named localization the LSP costs 6% to 118% more tokens and agents ignore it even when it is free, saving tokens only for the weakest model on reference-completeness tasks. On multi-file renames scored by real test execution, grep succeeds perfectly while a location-only LSP fails three-quarters of them, since a rename must touch comments and strings that semantic references exclude — pointing to an adaptive router keyed on task class, model capability and lexical noise rather than LSP-always.

Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems

Heming Fu, Shan Lin, Qianqian Xie, Guojun Xiong When an agentic system retries after a failed answer, the extra tokens open a gap between a model's advertised per-token price and what the workflow really costs — token inflation, defined as the ratio of true workflow cost to single-call cost and measured as high as 4.25× for a 7B model on multi-hop question answering. InflationAgent routes on that predicted true cost rather than list price, using CoT Branching Entropy, a pre-execution difficulty signal computed from local inference alone (AUROC 0.887), and selecting models by expected accuracy divided by predicted cost, with a policy that discards failed chains before escalating. On GSM8K under a fixed budget it reaches 94.7% accuracy versus 91.0% for FrugalGPT while using 31% fewer tokens, and forwarding a failed reasoning chain to GPT-4o is shown to cost up to 34.8 percentage points of accuracy, supporting the fresh-escalation design.

Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents

Bo Jin, Qiang Jiao, Xin Tong Tool-using agents that modify local state, keep persistent memory and speak external protocols bring risks of over-privileged actions, weak auditability, prompt injection, tool poisoning and uncontrolled side effects. Agentao is a local-first runtime that separates model-generated action proposals from host-authorized execution, layering host-facing surfaces, a host contract, a runtime core, a permission-mediated tool system, and subsystems for memory, replay, plugins, skills, sub-agents and protocol integration. The authors are explicit that no formal safety guarantees follow; the contribution is making permissions, state, protocol boundaries and execution traces explicit runtime abstractions, with the threat model, governance model and structured event interface described and the code released publicly.

Measuring Cross-Task Behavioral Consistency in Language Model Agents

Amritesh Banerjee, Pranil Raichura Agent evaluations report success rates, which say whether a system solved a task but nothing about whether it approaches different tasks the same way. The Behavioral Consistency Metric (BCM) trains a model to predict task success from behavioral features of execution traces, extracts a per-trajectory feature-attribution vector, and averages pairwise similarity of those vectors within an agent system. Over roughly 9,000 trajectories from six language model agents on software engineering tasks, within-task reproducibility and cross-task consistency turned out to be separate axes: some systems repeat themselves faithfully on retries of one task yet have no stable strategy across tasks, a split invisible to prior same-task reproducibility measures, and consistency did not track success rate.

MobileMem: Learning from a Year of Mobile Experiences

Xinle Deng, Yida Xue, Xiangyuan Ru, Haoming Xu, Shuofei Qiao, Mengru Wang et al. Personal assistants that accumulate knowledge about a user over months need benchmarks built from heterogeneous, multimodal, evolving personal data, which existing long-term memory evaluations do not provide. MobileMem supplies both a benchmark and framework for on-device memory, using a knowledge-grounded synthesis pipeline to turn a year-scale collection of real user-app sessions into coherent, temporally consistent long-horizon trajectories. It offers matched text-only and multimodal settings that probe multi-hop and temporal reasoning, updating superseded knowledge, and inferring preferences the user never stated, framing memory as accumulated experience rather than a retrieval index over isolated facts.

Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis

Aryan Luthra, Kshitij Jain, Siddharth Arya, Bobby Filar, Anna Bertiger Systems that pair a language model with retrieval or memory so it improves from feedback without retraining are hard to evaluate in security operations, where benchmark labels are scarce, stale, and unrepresentative. The proposed alternative dispenses with labels entirely: a stronger teacher model supplies sparsely sampled corrections to a smaller student running the harness, and the harness is scored by how far the student converges toward the teacher over time. Across security tasks, model families, and harness designs, teacher-relative lift correlated with improvement measured against a held-out gold standard, while LLM-as-a-judge comparisons between similarly capable models produced no usable signal at all.

SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data

Bruno Santos Teixeira Natural-language front ends to enterprise databases have to turn vague requests into queries that are correct, policy-compliant, cheap, and repeatable, and SemPlan compares four architectures for doing so on a deterministic synthetic benchmark of 1,800 English and Brazilian Portuguese cases (1,200 frozen for scientific evaluation). The contenders are direct SQL generation, a bounded tool-agent, structured semantic-request generation followed by deterministic planning and execution, and a clarification-capable stateful variant of the latter. Across 4,800 records answer correctness stayed low for every architecture — 22.25%, 22.58%, 25.67%, and 24.25% respectively — with the deterministic planner most correct, direct SQL safest on policy compliance, and the clarification variant cheapest, supporting a trade-off reading in which added structure shifts failure modes rather than eliminating them.

Ontology-Grounded Project Memory for Coding Agents

James Adam As coding agents generate more of a project's code, the reasoning behind those changes becomes hard to track. MOOSEDev stores architectural decisions, lessons, constraints, and rationales as typed records in a knowledge graph with lifecycle status, provenance, and supersession links, exposed to agents through a Model Context Protocol (MCP) interface and queried by a neurosymbolic engine that treats the symbolic layer as the primary reasoning substrate. On a neutral public corpus of 835 records, it returned essentially the full expected answer set (0.98-1.00) on supersession, set-completeness, and negation questions, versus 6% to 27% for a production vector-memory baseline, while relevance recall and token cost were comparable between the two.

Second Thought: Reasoning in Parallel as LLM Agents Act and Observe

Zhensu Sun, Chengran Yang, Yunbo Lyu, Jieke Shi, David Lo Agents in the ReAct paradigm think only during the Thought phase and sit idle while an action is serialized and the environment responds, leaving a recurring reasoning idle window unused. Second Thought is a training-free inference framework that forks four auxiliary reasoning branches the moment a Thought phase ends, decodes them concurrently with the main loop, and merges their output back when the observation arrives, moving the extra reasoning off the sequential critical path. Across three agentic benchmarks and three reasoning models it lowered average turn count in all nine model-benchmark pairs and cut main-thread decoding by up to 43% with Pass@1 unchanged in seven pairs; against a compute-matched control that spends the same budget on the main thread, it achieved strictly higher Pass@1 with 1.3 to 3.2 times less sequential decoding.

CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain Multi-agent retrieval pipelines usually debate a finished draft rather than the individual assertions inside it, and treat every retrieved modality as equally trustworthy, so hallucinations introduced between agents go unnoticed until the final text. CLAIR-Fin splits each question into atomic claims tracked in a typed ledger, conditions how much a claim trusts each evidence modality on the claim's type, checks grounding at the hand-off from drafting to adversarial review, and routes contested claims into a debate whose depth scales with what the debate turns up. On BB-FinQA-X, a 500-question cross-modal set built from Bangladesh Bank annual reports, faithfulness rises from 0.780 to 0.889 against a single-pass retrieval-augmented generation baseline while the system abstains on 5.4% of questions with insufficient evidence, also beating HyDE and Graph-RAG (both at most 0.874).

Simulation-Aware In-Context Policy Improvement for LLM-Aided Analog Layout Refinement

Bingyang Liu, Ziming Wei, Xiaohan Gao, David Z. Pan Analog integrated-circuit layout still depends on experts hand-tuning optimization parameters between rounds of parasitic extraction and post-layout simulation, and Bayesian optimization needs hundreds to thousands of such evaluations to help, which is far too many at layout level. The proposed multi-agent framework performs in-context policy improvement: agents run an act-observe-reflect loop over compact structured representations of the layout, updating the parameters exposed by an analog layout generator between simulations. On real analog circuits it improves post-layout performance over both the generator's built-in heuristics and Bayesian-optimization tuning using only tens of post-layout simulations.

From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL

Wenyue Hua, Zachary Huang, Tyler Payne, Safoora Yousefi, Saleema Amershi, Asli Celikyilmaz Agents acting for a user often face a counterpart with opposing goals, and the agreeable dispositions of a helpful assistant make for a bad delegate: frontier models volunteer private information and concede early. SocialRL trains social reasoning directly in a 4B-parameter model across six negotiation and scheduling domains including Deal-or-No-Deal, CaSiNo, Craigslist, and a job-interview setting, with every policy evaluated on all six. In-domain training closes 73–122% of the gap from baseline to frontier, and consolidating the specialists through cascade reinforcement learning and multi-teacher on-policy distillation yields a single 4B model averaging 0.627 utility across all six environments, matching or beating GPT-4.1, GPT-5.1, and GPT-5.2. Transfer follows game structure — structurally paired games lift each other while isolated ones transfer nothing — and distilling theory-of-mind traces rather than actions alone helps everywhere, though only next-action prediction correlates with negotiation outcomes.

AdsWorldEngine: A Self-Evolving Conversational Advertising Agent through Orchestrator and Tool Coevolution

Simiao Zuo, Chenhui Xu, Yimeng Jia, Qiang Lou, Jian Jiao, Denis Charles cross-listed Ads inside multi-turn assistant conversations must infer commercial intent from the query, the assistant's reply, and dialogue history, and also judge whether an ad would help or annoy. AdsWorldEngine splits this across an Opportunity Gate that decides whether to show anything, an Orchestrator that generates intents, calls advertising tools, and assembles a top-three slate, and an Evaluator that scores delivered ads offline. The Orchestrator is trained by supervised fine-tuning plus agentic reinforcement learning, and its high- and low-reward rollouts then become preference data for training the tools themselves, creating a loop that improves tool use and the tools together; subjective calls are handled by judgment models distilled from human labels with reflection-filtered rationales and a cost-sensitive GRPO variant. Offline diversity rises 60% and relevance 80% over the production system, and an online A/B test shows 22% higher revenue per mille and 74% more ad coverage.

Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model

Stephanie Jarmak cross-listed Coding agents are benchmarked as models but shipped as systems whose reliability depends on the harness, execution environment, retrieval, memory and state handling, permissions, review interfaces, and resource allocation. This monograph synthesizes 164 scholarly works, 100 practitioner records, 29 benchmark records, and 17 operated-system case records into a framework that treats evaluation and operation as a dependency chain, arguing that many apparent model failures actually originate elsewhere in the system, and that gains at one layer often fail to show up in end-to-end outcomes. It contributes a versioned catalog of 206 reliability records, an evidence ledger, runnable evaluation and reliability protocols, and five reusable agent skills, while noting that the review is structured rather than exhaustive and that evidence strength varies by topic.

MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends

Chaoqun Zhan, Qiang Zhou, Guannan Li, Zhenqiang Huang, Qianjin Wang Most agent-memory benchmarks probe after-the-fact recall rather than whether stored memory actually carries an agent through interdependent multi-session tasks, which is what MemoryArena measures. Holding the agent framework, model alias, task samples, and scoring code fixed, the study swaps only the memory backend, comparing MemoryLake against Mem0, vector retrieval over text-embedding-3-small, and a long-context control across five domains. MemoryLake leads on mathematics, physics, and progressive retrieval and averages 20.5% success versus 13.6% for the best comparator, though every system scores zero on travel planning and near-zero on web shopping, and the authors stress that sample sizes are small, confidence intervals overlap, and no significance tests were run.

Agentic Transaction: Towards ACID-Compliant Agent Systems

Zhaoyan Sun, Xiaoxiao Wang, Guoliang Li cross-listed As language model agents move from chat into long-horizon work over persistent workspaces, they hit the same problems transactional databases were built to solve: reliable execution, consistent outcomes, safe concurrency, and durable state. The proposed framework reinterprets atomicity, consistency, isolation, and durability as semantic guarantees for agent execution, and instantiates them in a data agent using transactional exploration-execution-validation cycles, transactional skill hubs, confidence-divergence validation, semantic dependency-aware isolation, and transaction-aware state management. On widely used benchmarks the system reports a 10.6% improvement over state-of-the-art agents including Claude Code.

When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict

Lu Yang, Shusheng Xu, Zhuoran Li, Tongkai Yang, Longbo Huang Agents that carry personal memory across sessions inevitably accumulate contradictions, and when a query provides no context, timing, or source authority to resolve them, picking one memory as definitive turns genuine uncertainty into a confidently wrong action. TANGLE is a benchmark of 541 deliberately unresolvable conflicts across 40 personas, split into context-partitioned, behavior-oscillation, and source-contradiction types, scored on conflict perception, causal reasoning, confidence calibration, clarification seeking, and memory faithfulness under both a curated-memory oracle track and an end-to-end extraction pipeline. Models notice conflicts far more reliably than they calibrate their actions or ask targeted clarifying questions, and pipeline extraction discards the conflict-bearing relations downstream reasoning needs, motivating a policy that adapts the response to the specific conflict rather than applying fixed rules.

AI Research Preference Models

Thomas Simon Foster, Bassel Al Omari, Tingchen Fu, Thomas Mann, Carl Domond, Lucia Cipolina-Kun et al. Research agents can write candidate machine learning experiments in minutes but need hours to days of GPU time to evaluate each one, so progress depends on how a fixed execution budget is allocated across proposals. The approach builds preference models (RPMs) from frozen pretrained language models with no task-specific training, in two forms: an inference-only variant that reasons over candidate plans, code, and previously executed solutions, and an agentic variant that first runs small-scale pilot experiments. Wired into the AIRA-dojo search agent and measured on AIRS-Bench, the two variants lift the average normalized score from 0.684 to 0.711 and 0.729, and reach the unguided agent's 24-hour performance in roughly 15 hours using under two-thirds of its execution budget, with new state-of-the-art results on two benchmark tasks.

HELIX: Model-Harness Co-evolution for Recursive Self-Improvement

Tianyu Fan, Chao Huang An interactive agent acts through a runtime harness that governs context, tools, control flow, and stopping, which means the harness shapes both what a model can do and the trajectories it later learns from. HELIX is a source-traceable substrate that decomposes agent systems into typed ports, reusable atoms, recipes, product shells, and runtime policies, making each intervention explicit and auditable while retaining trajectories, test outcomes, and provenance, so harnesses can be evolved for a fixed model and then rebuilt as the model improves. In one evolution round on code repair evaluated with the SWE-bench evaluator, a 65-candidate portfolio found a fixed harness improving task coverage by 4.0% over Pi, while the full portfolio exposed up to 58.0% more verified coverage through complementary sibling behavior, and a 200-slot sibling slice yielded 438 verified supervised, critic, filter, and preference training records.

MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning

Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang et al. cross-listed Answering what happens before, after, or across stages of a tens-of-minutes surgical video requires grounding a question in evidence spread across time, which one-shot vision-language models lose by compressing the procedure into a context window, while trained video agents are data-hungry and transfer poorly to unseen surgery. MedClaw separates reasoning from perception: a text-only orchestrator plans what evidence to gather and issues an auditable sequence of tool calls that frozen vision-language sub-agents execute over pixels by viewing, cropping, inspecting frames, and retrieving external knowledge, and a gradient-free, reward-gated Heuristic Skill Distillation loop mines low-scoring traces and keeps a candidate skill only if it raises validation reward, yielding reusable retrieval skills such as directed re-look. Because it grows an external skill library instead of tuning weights, the loop adapts from roughly 100 labeled examples, and on the new MedClawBench of 1,123 doctor-grounded questions over long neurosurgery recordings plus a held-out lecture-video split it beats one-shot models and general video-agent frameworks across all four evaluation dimensions.

Agent-Orchestration in Autonomous Chip Design

Linyang Li Framed as a position piece on where tool-using language-model agents fit in integrated circuit design, the work proposes treating a chip-design system not as a single model but as a large AI-organization of coordinated agents. The argument centers on what kind of artificial intelligence the industry actually needs given how specialized and interdependent the design flow is. The contribution is the organizational framing itself rather than a benchmarked system.

Demystifying Agent Skills: Why They Work-Until They Don't

Zhiyuan Jiang, Fangrui Huang, Hanwen Xing, Xander Wu, Yipeng Gao, Rui Cao et al. Agent skills — structured packages of procedural knowledge injected at inference time — are widely reported to raise task success, but little is known about the mechanism or the failure modes. Controlled experiments across several benchmarks, agent harnesses, and language models isolate the effects of skill representation, outcome annotation, retrieval difficulty, and cross-framework robustness, and a paired-trajectory contrastive study normalizes 8,135 trial records into a taxonomy of twelve skill-use modes. Procedural anchoring — stabilizing the sequence of actions rather than supplying missing facts — accounts for 65.7% of cases where skills help, versus 4.5% for explicit knowledge injection, and skills beat workflow memory by 6.06 points in matched comparisons. Retrieval turns out to be a separate bottleneck: as the skill pool grows from 5 to 100, the precision of skills actually used falls from 29.6% to 3.3%, though exact ground-truth skill selection proves neither necessary nor sufficient for downstream success.

MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation

Juli Huang, Hannah Clay, Sajjad Beygi, Thomas Sarda, Negin Golrezaei, Amin Saberi cross-listed Merchant-deployed conversational shopping assistants must recommend only from a fixed catalog, honor hard constraints like budget and brand exclusions, and keep those constraints straight across turns — reliability requirements that prompt-only language-model systems handle poorly. MACS splits the work: the language model interprets requests, elicits preferences, and writes responses, while product retrieval, hard-constraint filtering, brand exclusion, and progressive constraint relaxation run deterministically in a merchant agent backed by a session-persistent preference layer. On a 140-query single-turn benchmark it reaches an 87.1% pass rate with perfect brand compliance, and on a 10-scenario multi-turn benchmark 72% macro Pass@5 with zero constraint drift versus 56% and 52% for catalog-bound GPT and Gemini prompting, with the largest margins on exclusion reversal and accumulated constraints.

A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents

Ismail El Hamraoui, Sagar Jose, Nicolas Bureau, Robert Plana Long-running LLM agents can silently drift away from their assigned task and cause irreversible side effects on external systems, and prompt-level guardrails offer no step-level detection, risk scoring, or recovery decision. Because the main executor is usually a large model that cannot be retrained per deployment, this work trains a separate small language model with reinforcement learning to occupy each node of an external recovery graph — drift classification, operation detection, risk evaluation, final decision — emitting XML-structured reasoning tailored to that role, with rewards combining schema and length rules with an LLM-as-judge score for semantic quality. On the public AppWorld benchmark the small model generally makes correct recovery decisions when given information about the suspected drift onset, and reliably respects the prescribed output schema at every node.

Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions

Xiaokai Yan, Jingtao Ding, Yong Li, Zhiwen Yu cross-listed Mobile GUI agents built on multimodal large language models (MLLMs) mostly react to explicit instructions, with no pipeline for inferring what a user is trying to do, predicting the next intention, and acting on it. Act2Intention supplies both a benchmark — 72,511 intentions and over 700,000 actions across 52 apps, collected and validated as continuous intention-action trajectories — and an agent that chains proactive intention understanding, personalized prediction, and experience-guided execution. Supervised fine-tuning on the benchmark yields absolute gains of +32.0, +10.25, and +6.9 points on understanding accuracy, prediction accuracy, and step success rate over the same framework without fine-tuning.

AutoSchema: Live Schema Grounding for Agentic Text-to-Sparql over Heterogeneous Knowledge Graphs

Yiming Zhang, Koji Tsuda Life science knowledge graphs each expose their own SPARQL schema and identifiers, so agents querying them typically depend on hand-curated metadata files that require language-model drafting plus manual review to maintain. AutoSchema instead does live schema grounding with no training: it inspects endpoints directly, maps entity names in a question to graph identifiers, explores relation paths, and discovers cross-resource links during iterative query construction. Against TogoMCP as the curated-file baseline, it raises mean factoid accuracy on biomedical knowledge-graph question answering while using fewer tool calls and exhausting its iteration budget less often, with consistent gains on a longitudinal BioASQ Task B evaluation and preliminary evidence of transfer to an undocumented chemistry RDF graph.

Polaris : Multi Agentic System for Conversational Enterprise Analytics

Varuni H K, Soham Sarkar, Jay Kumar, Goutham Krishnan, Tanvi Johari, Avinash Bharadwaj et al. Enterprise analytics stalls when business questions require chaining query generation, visualization, and explanation across systems that non-specialists cannot address directly. Polaris is a supervisor-led multi-agent framework whose Dynamic Task Coordination layer treats agent-task assignment as adaptive bipartite matching, letting the supervisor reassign and recover across specialized querying, visualization, and reasoning agents at runtime; these agents follow a reason-first ReAct loop so answers include the underlying explanation rather than just retrieved rows. Evaluation on structured enterprise datasets reports high semantic fidelity and answer relevancy, though the paper describes results qualitatively rather than against a public benchmark.

TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments

Qingren Yao, Yaxuan Kong, Yuqi Nie, Yichen Li, Stefan Zohren, Anna Vettoruzzo et al. Time series question-answering benchmarks are built from frozen snapshots, so they never test whether an agent respects data cutoffs or revises conclusions when a new release changes the evidence. TimeSage-EV tracks 60 real institutional scenarios across 6 domains with 1,485 scenario-period question-answer pairs spanning February 2023 to May 2026 and monthly, weekly, daily, and irregular release cadences; at each period an agent sees the series and source reports while the withheld next release supplies ground truth for state identification, summarization, and outlook reasoning. Experiments with frontier LLM agents and TimeSage-1.0, a self-evolving agent that accumulates a reusable analytical skill library, show large gaps between model tiers and recurring failures in temporal validity, use of exogenous context, and adaptation to new releases. The benchmark is released with monthly updates, code, a leaderboard, and failure-mode analyses.

Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents

Zhizhao Guan, Chen Huang, Ziming Liu, Hongru Liang, Wenqiang Lei, See-Kiong Ng et al. Agents built on language models tend to act on what they already know rather than deliberately gathering information that would improve later decisions, a capability the authors call proactive exploration. SAFARI attacks two identified bottlenecks: standard expert demonstrations suffer hindsight bias because the demonstrator already knew the answer, so an Exploratory Data Construction stage synthesizes trajectories rich in genuine information-seeking; and reward signals conflate useful exploration with aimless wandering, so reinforcement learning is guided by contrastive trajectory pairs that separate productive exploration from redundant wandering. Experiments across environments report gains from both components along with analysis of what proactive exploration looks like in practice, and the code is released publicly.

ATLAS: Discovering Agent Strategies through LLM-Guided Abstraction and Automata Learning

Ignacio D. Lopez-Miguel, Andreas Happe, J\"urgen Cito, Ezio Bartocci, Bettina K\"onighofer, Martin Tappler cross-listed Agents driven by language models are now run on software testing and cybersecurity tasks, but evaluation usually stops at task success and raw execution traces, which say little about the strategy an agent actually followed. ATLAS (Automata Learning for Agent Trajectory Analysis and Strategy Discovery) abstracts trajectories into symbolic events and applies automata learning to infer finite-state models of agent-environment interaction that expose recurring behaviours, decision points, successful completion paths, and failure loops. Applied to a penetration-testing agent across 12 vulnerable machines, the learned automata surface high-level exploitation strategies that are hard to read off raw traces, and the authors further demonstrate transferring the extracted symbolic model from a frontier model to a compact one, plus model transformations that yield concise behavioural explanations.

ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond

Mingming Zhao, Jiqian Dong, Kangping Xu, Zadid Hasan, Chengrui Fan, Shan Jiang et al. Autonomous research agents stall over long horizons because they lack ways to preserve continuity, back out of dead ends, and spend compute where it is paying off. ScienceFlow structures a research run as segments grounded in executable workspaces, representing progress as recoverable executable states; transitions are governed by Executable-State Transition through Re-Anchoring (ESTRA), which picks either the live state or an archived one as the next anchor and decides whether to continue or redirect, while an evidence-aware controller allocates compute to jobs based on availability, remaining budget, and validated progress. Across machine learning, scientific modelling, and mathematical optimization tasks, it reaches 70.22% Any-Medal on the full MLE-bench within a 24-hour budget, 4.92 percentage points above the prior best reported result.

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang, Weijian Zheng, Fernando Llorente, Xiaolong Ma et al. Multi-agent reasoning pipelines usually decide which intermediate messages to keep by proxies for correctness such as agreement, confidence, or automated scoring, which assumes a wrong message is not worth passing on. Diverse Hypothesis Deliberation tests that assumption by caching five independently generated messages and replaying the same downstream solver twice, once with a given message visible and once with it hidden, so the causal effect of each message on final correctness can be measured directly. Across five mathematics and science benchmarks with gpt-oss-120b and gemma-4-31B-it, messages carrying wrong answers but useful decompositions or constraints appear in every benchmark-model pairing, and more than four in ten of the correctness-flipping wrong-answer messages flip it in the helpful direction. Passing the complete message beat passing only its reasoning, which in turn beat passing only its answer, and the resulting replay labels can be reused to train when agents should listen to each other.

AgentRewind: Recoverable Execution for Long-Horizon LLM Agents

Yu Zhuang, Kefei Chen, Yitong Duan, Shuxin Zheng, Jian Li, Xu-Yao Zhang When a language-model agent works over a long horizon, an early mistake contaminates both its context and the environment state, and later actions often cannot undo it; existing defenses concentrate on preventing errors through planning and safety checks rather than recovering from them. AgentRewind is a runtime framework that records aligned checkpoints of the agent's context and a controlled environment, letting the agent roll back to an earlier state and retry while carrying forward what it learned from the failed attempt. The authors also build MettleBench, a benchmark of long-horizon engineering assignments made of chained requirements that scores both completion and partial checklist progress. Across several models, execution strategies, and agent harnesses, rollback-and-resume raises both task success rate and average checklist progress over the baselines.

The Past and Future of AI Scientists

Ross D. King Machines that originate hypotheses, deduce consequences, design and run experiments, and revise beliefs have existed since Adam made the first novel machine discovery through physical experimentation and Eve established the self-driving-laboratory architecture; foundation models, autonomous agents, and lab robotics now make far more general systems buildable. Surveying that history and what comes next, the authors argue the open problem is no longer automating individual components of science, which is already possible, but integrating them — combining neural learning with logic, probability, mathematics, causal reasoning, simulation, experimental design, robotics, and formal scientific records. They assess progress against the Nobel Turing Challenge's 2050 target for automating Nobel-quality discovery and judge it ahead of schedule.

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Yuhao Zhan, Bingxiang He, Zecong Tang, Chaojun Xiao Self-evolving agents are usually measured under fixed execution conditions, so nothing tests whether they can recover once the environment changes underneath them. PACE-Bench supplies 144 source-to-target adaptation pairs across six simulated physics domains, where a code-driven design that works in the source environment breaks in a mutated target with the same goal and interface, and the agent must repair it using sandbox diagnostics within a limited attempt budget. Across ten methods from four paradigms the benchmark stays far from saturated — Reflexion with Qwen3-14B solves only 35.9% of pairs, and GPT-5.5 reaches 66.7% on the Statics subset alone. Simulator-grounded reflection beats unverified self-revision, memory tends to anchor agents to their initial design, and even disclosing the exact physical change does not lift the ceiling, suggesting the bottleneck is redesigning mechanisms rather than inferring parameters.

Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports

Beatrice Alessandra Motetti, Emilien Guandalino, Daniele Jahier Pagliari, Alessio Burrello, Lorenz K. M\"uller, Konstantin Berestizshevsky et al. Generative models used to synthesize technical write-ups tend to produce ungrounded claims, and they rarely handle non-text content well. Wyvern is a multi-agent pipeline that assembles technical reports combining text, figures, and tables with supporting citations, adding an automatic claim-revision stage aimed at keeping assertions tied to sources. In a human study, evaluators judged its figures more informative than a recent baseline in 87% of comparisons and preferred its reports over three alternatives in 63% to 100% of cases, while automatic scoring showed up to 2.3× higher citation recall and 1.6× higher citation precision than the baselines.

SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning

Panjing He, Mingyue Cheng, Yucong Luo, Li Li, Xiaohan Zhang Real workbooks carry implicit cross-table associations, fine-grained column dependencies, and spatial layout that large language models lose when a sheet is flattened into a linear string. SheetCompass instead builds explicit hierarchical relation graphs capturing structure within and across worksheets, and pairs them with a memory component that retains task-relevant state as an agent works. The claim is that preserving intra-sheet boundaries and inter-sheet semantics lets agents use the global spatial context human analysts rely on when reasoning over and automating complex spreadsheets.

Twin: Playing an Unknown Game with a Test-Time Digital Twin

Alexy Skoutnev, Kirill Acharya, Gaston Longhitano, Madeleine Udell, Kevin Ellis, Iddo Drori Playing a game whose rules and goal are hidden usually means hand-engineering a world model per task; Twin instead has a frontier coding agent write an executable world model from interaction alone, on continual-learning tasks such as ARC-AGI-3. A harness enforces replay validation — no action is taken until the program reproduces every previously observed transition — and each prediction mismatch becomes a counterexample that drives a repair of the program. The system clears 179 of 183 levels (97.8%), beats human efficiency on 158 of them, and infers the goal before receiving any reward on 156; on the benchmark's human-referenced 0–100 score the same base model goes from 7.8% played directly and 61.1% with an off-the-shelf harness to 93.3% with the twin world model. The authors conclude that constructing a usable world model was easier than expected, while identifying the right goal is the harder half.

Other 29

The Architect: Interactive Visualization of Deep Learning Mathematics Directly in Microsoft Excel

Mohammad Imrul Jubair, Tom Yeh cross-listed Deep learning libraries hide the arithmetic behind function calls, and most visualization tools stop at architecture diagrams or training summaries. The Architect takes a compact table describing a neural network and generates a Microsoft Excel workbook in which the full forward pass — and, on request, backpropagation and parameter updates — appears as live spreadsheet formulas that recompute automatically when the user edits inputs, weights, labels or hyperparameters. Matrices, activations, losses and gradients become inspectable spreadsheet regions, aligned PyTorch snippets connect formulas to code, and the report walks through arithmetic tracing, learning-rate exploration, and diagnosis of dying ReLU units and vanishing gradients.

AI Evaluation Should Work With Humans

Jan Kulveit, Gavin Leech, Tom\'a\v{s} Gaven\v{c}iak, Raymond Douglas A position argument that the prevailing evaluation paradigm — measuring superhuman autonomous performance — implicitly aims AI development at replacing people rather than complementing them, and is steering the field in the wrong direction. The proposed pivot is to evaluate human-AI teams rather than models in isolation, on the grounds that team-level metrics would push development toward systems that genuinely complement human capabilities and produce better societal outcomes.

Hard Cases, Bad Labels: Testing Error Exposure and Error Location in Uncertainty Sampling Under Bounded Label Noise

John Myron Uy Active learning picks the examples a model is least sure about, but those same examples tend to be the ones human annotators get wrong, raising the question of whether uncertainty sampling suffers because it collects more bad labels or because bad labels in hard regions hurt disproportionately. Margin-based selection is compared against random sampling on three public binary tabular datasets under clean labels, random classification noise, and bounded difficulty-dependent noise, using 100 paired seeds, nine noise rates, and budgets from 20 to 120 labels with regularization re-tuned at each budget. Uncertainty sampling gained 1.09 to 1.77 percentage points of normalized balanced-accuracy area under the learning curve with clean labels, and exposure-matched controls found no evidence for a universal extra penalty from errors being concentrated in difficult regions — the apparent robustness varied by dataset, budget, noise structure, and which metric was used.

Asymmetric Discourse Homogenization and Shared Language Technology: Evidence from Reddit

Fengming Liu cross-listed Using six million Reddit comments from two cross-partisan forums between 2019 and 2025, the analysis finds an ideologically asymmetric break around late 2022: conservative users' prior trend toward more diverse political discourse was interrupted while progressive users showed no comparable change, a pattern that holds across interrupted time series, difference-in-differences, regression discontinuity in time, and propensity-score matching. A permutation test over 2,377 candidate cutoff dates places the ChatGPT release at an unremarkable 49.8th percentile, indicating gradual buildup rather than a single break, though a continuous cumulative index of exposure across seven model releases stays significant under a quadratic trend. Restricting to authors active throughout the window makes the effect vanish, so the homogenization looks community-level and ecological rather than driven by individual users adopting AI writing tools, with concurrent secular change not fully ruled out.

Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection

Tom\'as Andrade Weber cross-listed Physiology constrains how human speech unfolds over time, suggesting synthetic speech should violate those temporal regularities in detectable ways. A causal LSTM next-frame predictor trained only on bonafide speech over features from the Wav2Vec2-Large-AntiDeepfake backbone tests this directly, with a global-average-pooling baseline on identical features isolating the contribution of temporal modeling, plus an optional supervised stage that fits a multilayer perceptron on the frozen LSTM states. Across ASVspoof 2019/2021, Codecfake, In-the-Wild, MLAAD-EN, and Deepfake-Eval-2024, the system reaches a best published 0.75% equal error rate on ASVspoof 2021, and the bonafide-only stage alone beats the published supervised baseline from the same backbone on Deepfake-Eval-2024. Static and dynamic features tie on near-domain data, while trajectory dynamics provide the gains on harder cross-corpus benchmarks.

Emergent Models: Intelligence from Tiny Substrates

Giacomo Bocchese, Nicola Giacobbo, Etienne Guichard, James Wiles, Akshaj Devireddy cross-listed Emergent Models treat modeling not as learning a closed-form input-output map but as coaxing computational behavior out of simple open-ended substrates such as cellular automata, which iterate a fixed local rule over a latent space for an adaptive number of steps with an interface connecting latent state to external signals, trained by evolutionary search. The hypothesis is that some such systems are biased toward global generalization, capturing the rule that generated the data over its full domain and extrapolating past the training range. Theoretically, certain instances are proven latent-universal: with the update rule and interface fixed, varying only the initial latent state realizes any partial computable function; empirically, a zoo of minimal discrete and continuous substrates with tens to hundreds of parameters extrapolates exactly on simple arithmetic and supports control and online adaptation, while exposing clear limitations. The stated aim is to widen the design space beyond differentiable feed-forward maps rather than to offer a competitive architecture.

Attributing Preprocessing Invariance in Spectral Foundation Models

Dongjun Wei, Hongyi Wu, Yinuo Zou Spectral foundation models are often credited with learning invariance to preprocessing because a classifier trained under one pipeline still works under another, but these models normalize inputs before any learned weight is applied. Analyzing a Raman spectroscopy foundation model, the authors show that per-spectrum normalization collapses two differently preprocessed spectra to the same vector exactly when one is a positive multiple of the other plus a constant — a form many standard preprocessing steps take — so the encoder never sees a difference to be invariant to. Measured against its own parameter-free normalization on six Raman datasets, the model shows no measurable gain, and a controlled experiment confirms it only learns to ignore a transformation that actually reaches it. A numerical test identifies which transformations a given normalization already removes; surveying released systems across five modalities, most normalizations remove such transformations, and replication on two systems claiming learned invariance again found no gain.
22 more specialized papers

Theory 20

Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning

Yongchao Huang A hidden Markov model (HMM) does three things — infer a belief over hidden state, propagate it through a transition, and emit back into observation space — and the argument here is that time-indexed Predictive Information Bottleneck VJEPA (PIB-VJEPA) has exactly that structure, with the stochastic context encoder acting as an amortized filter, the probabilistic predictor as latent dynamics, and the decoder or inverse target encoder as the emission. Four progressively stronger levels of correspondence are defined along with sufficient conditions for exact sequence-level HMM equivalence, and MCJEPA makes the link concrete by replacing the latent predictor with a learned transition matrix whose powers guarantee exact multi-horizon Chapman-Kolmogorov consistency in the finite time-homogeneous case. Deterministic temporal JEPA falls out as a degenerate Dirac-kernel case, and controlled experiments confirm transition composition, the filtering interpretation, and predictive Markovization on a synthetic process with known structure.

Consistent Model Chasing Is Minimax Optimal: The Exact Value of Scalar Adversarial Adaptive Control under Large Parametric Uncertainty

Dimitar Ho cross-listed For the scalar system x_{t+1} = a x_t + u_t + w_t with bounded disturbance and an unknown pole of unknown sign and arbitrarily large magnitude Δ, prior adaptive control theory offered stability certificates, gain bounds, and regret rates but never the exact minimax peak deviation. The analysis pins that game value at **γ*(Δ) = 1 + Δ**, splitting it into the irreducible cost of the disturbance plus the unavoidable price of a single identification spike, and shows the optimal policy is certainty-equivalent deadbeat control at the midpoint of the set-membership consistent interval. Writing the controller as an oracle-selector composition shows this architecture is forced rather than merely sufficient, and the optimal law contains no exploration at all: probing is punished before it pays, commitment is fatal under weak excitation, and optimism costs asymptotically at least twice the optimum.

On the Brittleness of Maximum Likelihood Estimation for Gaussian Process Hyperparameter Optimization

Tyler R. Johnson, Kian Ben-Jacob, Christopher P. Muller, Ramin Bostanabad cross-listed Maximum likelihood estimation (MLE) is the default way to fit Gaussian process (GP) hyperparameters in engineering design, and GPs are widely assumed to resist overfitting, but MLE degrades badly when its assumptions fail. Theoretically grounded alternatives to the MLE objective are compared against it for probabilistic regression and classification, along with practical recipes for choosing among them. In downstream Bayesian optimization tasks the resulting GPs beat tabular foundation models such as TabPFN on prediction accuracy, uncertainty quantification, and inference cost, giving practitioners a concrete blueprint rather than a default loss.

What preferences can - and cannot - predict in multi-agent online learning

Omar Abbadi, Rida Laraki, Panayotis Mertikopoulos cross-listed Solution concepts in game theory are ordinal — defined by which outcomes players prefer — while no-regret learning dynamics such as follow-the-regularized-leader are continuous-time processes, and it is unclear how much the first determines the second. One direction is settled cleanly: the pure profiles contained in any dynamically stable set must be closed under profitable deviations, and for subgames obtained by restricting action sets, preferences fully characterize asymptotic stability. The converse fails in general — a three-player game is constructed with a preferentially stable set whose span is dynamically unstable — and the gap is bridged by resilience under aggregate deviations, an easy-to-check payoff condition sufficient for asymptotic stability of arbitrary spans of pure strategies.

When Does More Correct Data Hurt? Insertion-Stability and the Limits of Dimension-Based Theory

Joseph Sankoorikal Johny Adding more correctly labeled data ought to be harmless, but under a monotone adversary that reads an i.i.d. sample and appends any examples the target hypothesis labels correctly, classes of VC dimension at least 2 provably cost a logarithmic factor above the clean PAC rate. Since that bound is a worst case, this work asks which classes actually pay, and answers that it depends on the learner: a learner is called insertion-stable if more correct examples can only shrink its error region, and such learners are immune because risk after insertions never exceeds risk on the clean part alone. Because the Closure algorithm is insertion-stable, every intersection-closed class keeps its clean rate, while classical dimensions cannot predict immunity — two classes can share VC and Littlestone dimension 2 yet split between the clean and penalized rates, and intervals have unbounded Littlestone dimension but are immune; on the known hard class, no monotone permutation-invariant compression scheme of any finite size attains the clean rate.

Boosting Data Augmentation with Stochastic Weight Averaging

Longde Huang, Axel Flinth, Jan E. Gerken Data augmentation is the usual way to bake task symmetries into an ordinary network, and recent theory shows infinitely large deep ensembles trained on augmented data become perfectly symmetric — but ensembles require repeating training many times. Stochastic weight averaging is studied as a cheaper substitute that needs only one run, analysed by approximating the late-training stochastic trajectory with an Ornstein–Uhlenbeck process. In the infinite-width limit, averaging weights over augmented training yields an equivariance gain larger than what the accompanying accuracy improvement alone would predict, a claim backed by experiments spanning computer vision and graph classification under both discrete and continuous symmetries.

AI-Assisted Discovery and Construction of a Counterexample to the Convergence of Three-Block ADMM with the Identity Matrix as its Third Constraint Block

Kenan Xu, Xiangfeng Wang cross-listed The alternating direction method of multipliers (ADMM) converges for two blocks but can diverge with three, and one subclass had stayed open: whether divergence is still possible when the third constraint block is the identity matrix. Using Codex driving GPT-5.6 Sol, the authors construct an explicit rational instance with strongly convex quadratic first two blocks and verify it along a piecewise-affine reduction path, showing by exact arithmetic that direct three-block ADMM enters a bounded nonconvergent orbit of period 66. A follow-up study of multiplier relaxation shows a problem-dependent small dual step can restore convergence on a fixed instance while no positive relative step works uniformly over the class, and an independent run with Kimi Code on Kimi K3 produced a different certificate (a locally attracting period-23 orbit), suggesting the research harness itself shapes which mathematical objects get found.

LLMs Don't Pay for the Jump

Paras Balani, Subhrakanta Panda Responding to an argument that language models cannot make the abductive leap that produced Einstein's equivalence principle because they lack embodied simulation, this position piece proposes a different missing ingredient using Planck's 1900 quantization of blackbody radiation as the test case. Planck's postulate needed no sensorimotor grounding — it was forced by a mathematical consequence of classical theory, an infinite predicted energy for a finite measured quantity — and the authors argue neither induction nor deduction could have produced it. They formalize the requirement as a thermodynamic coupling in which epistemic error carries physical cost, and argue fixed-weight transformer inference has no such coupling at any scale, consistent with the empirical observation that output entropy barely changes as accuracy on increasingly hard causal tasks falls from 100% to 17%.

The Dynamics of Intelligence Explosions

Toby Ord If AI systems increasingly do AI research and development, the resulting feedback loop might in principle produce runaway capability growth, and this analysis works through the mathematics of the most explosive cases to see what actually drives them. Singular growth toward a vertical asymptote turns out to be harder to reach than recent economics-inspired models suggest, and there is a neglected middle class of trajectories that grow faster than exponentially yet never hit an asymptote. The pivotal and largely overlooked parameter is generation time, the time to go once around the feedback loop: singular growth is impossible unless generation time rapidly approaches zero.
11 more specialized papers

Reinforcement Learning 16

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

Yongqi Tong, Tan Li Hui Faith, Choy Zhen Wen Marcus, Zhou Jin, Kewei Fu, Jiang-Ming Yang et al. Group-based reinforcement learning assumes the rollouts compared within a group are doing the same kind of thing, but in open-ended interaction an agent might answer immediately, ask a clarifying question, post a progress update, or confirm before acting — all valid, and a reward model's stylistic preferences then contaminate the relative advantages. Framing this as a reward fairness problem, ARC (Advantage Regularization via Conditioning) groups rollouts by interaction strategy before computing advantages, combined with hybrid rewards and entropy regularization, and is studied inside an interaction paradigm that separates what the user sees from latent reasoning and tool calls, supported by an 86K-example strategy-annotated training corpus. ARC substantially improves scores on the tau and tau^2 tool-use benchmarks, while the decoupled interaction design cuts time-to-first-token from 4.91 to 1.27 seconds versus a think-first baseline.

Reward Machines for Signal Temporal Logic

Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai Signal temporal logic (STL) specifies real-time properties over real-valued signals and supplies a quantitative robustness score, which prior work has fed directly to reinforcement learning as a reward; the trouble is that robustness depends on the entire execution history, so the state space blows up for long-horizon specifications with nested temporal operators. The proposed approach compiles a specification into a timed alternating automaton, augments the state space with automaton locations and clock valuations to serve as a compact memory, and derives Markovian rewards from the automaton's acceptance condition. Policies trained this way achieve higher robustness scores and satisfaction rates than those learned from robustness-based rewards.

GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

Konstantin Dobler, Federico Scozzafava, Jonathan Janke, Mohamed Ali, Simon Lehnerer Reinforcement Learning with Verifiable Rewards (RLVR), typically optimized with Group Relative Policy Optimization (GRPO), is the standard recipe for improving reasoning in language models, but nearly all published studies train and evaluate in English. A large-scale sweep varies the base model, the training language, and the reward applied to the language the model reasons in. Training a model to reason in its native language costs only a small amount relative to reasoning in English, and training in a single language often transfers gains to many other languages. The effects are strongly model- and language-specific, though, with some training languages causing severe regressions on out-of-domain capabilities elsewhere, so multilingual RLVR needs broad evaluation to catch them.

Dynamic Multi-Depot Vehicle Routing with Online Requests: Event-Driven Transformer--DRL and Rolling-Horizon Benchmarking

Faezeh Ardali, Gerald M. Knapp Dispatching vehicles from multiple depots when requests arrive over time requires decisions in milliseconds, making it a natural target for learned policies. Masked multilayer perceptron and Transformer policies were trained by behavior cloning and proximal policy optimization (PPO) with deterministic feasibility masking and fixed-prefix route commitments, then compared against insertion heuristics and time-limited rolling-horizon optimization on a shared 20-scenario protocol. Every method served all requests without invalid actions, but a simple nearest-feasible heuristic achieved the lowest objective and beat both learned policies on routing quality, waiting time, stability, makespan, and runtime; rolling-horizon optimization won on waiting time and makespan at far higher compute. The learned policies did keep millisecond inference and transferred to 80-request instances without retraining, and PPO helped the Transformer on average while adding seed variance.

Learning to Run Power Networks: Effective AlphaZero-inspired Topological Control

Lukas Zetto, Benjamin Sch\"afer, Qiong Huang Reconfiguring grid topology is a cheaper way to relieve congestion than redispatching generation, but the action space is combinatorially huge and operational constraints are strict. This study applies AlphaZero-style planning with Monte Carlo Tree Search (MCTS) to proactive grid operation and systematically varies reward design, observation density, and search guidance. The tuned agent reaches 98.43% peak survivability, well above a proximal policy optimization (PPO) baseline, and — counterintuitively — running MCTS without a learned prior policy or value function improved training efficiency, while a plain binary survival reward guided search better than multi-objective alternatives. The authors conclude that pure reinforcement learning is insufficient and that domain heuristics, binary rewards, and a restricted line-load observation space are what make the system work.

Reinforcement Learning-Based Production Scheduling in an Industry-Based Coating Scenario Using the Digital Model Playground

Arne Kr\"oger, Ralf Buscherm\"ohle, Wilhelm Hasselbring, Henrik Wilbers Reinforcement learning results for production scheduling mostly come from simplified benchmark shop problems, which limits what factories can conclude from them. This work models an industry-inspired coating process with sequence-dependent setup times, machine breakdowns, due dates, and variable utilization in the open-source Digital Model Playground (DMPG) discrete-event simulation framework, then trains Deep Q-Networks and Proximal Policy Optimization agents against conventional dispatching rules. Both agents improve key performance indicators in a balanced way, with PPO giving the most robust performance, and the shareable scenario plus framework is offered as a reusable testbed.

Traj-LeWM: Path-Aware World-Model Planning via Latent Trajectory Cost

Xiaodi Huang, Ziyi Ding, Jingtian Wan, Yuchen Liu, Yuan Zhang, Xiao-Ping Zhang et al. LeWM is a pixel-based latent world model that scores candidate action sequences purely by how close their predicted endpoint lands to the goal, so it learns only local next-step transitions and ignores what the rest of the predicted path looks like — a problem when predictions diverge from execution. Traj-LeWM keeps that objective and endpoint score but adds a goal-conditioned latent trajectory cost that summarizes the whole predicted rollout, used during training as trajectory-preference supervision and during planning as a second ranking signal. It beats LeWM by 3, 14, 7, and 7 percentage points on Push-T, OGBench-Cube, Reacher, and Two-Room, with ablations separating the representation-shaping and candidate-ranking contributions.

Deep Reinforcement Learning solution for pickup and delivery routing problems with time window and capacity constraints

Andrew Soroka, Alex Meshcheryakov, Sergey Gerasimov Real-time vehicle routing for goods pickup and delivery becomes intractable for classical methods once capacity and time-window constraints are added at medium to large scale. A modified version of the JAMPR deep reinforcement learning model is applied to the Pickup and Delivery Problem with Capacity and Time Window constraints (CPDPTW), reportedly the first successful deep RL solution for this constraint combination. The learned policy produces fast optimal solutions on small and medium instances and fast suboptimal ones beyond roughly 200 nodes.

Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine

Chenran Weng, Joo Seung Lee, Malini Mahendra, Anil Aswani Reinforcement learning for mechanical ventilation settings usually consumes only structured electronic health record data, discarding the clinical context that lives in free-text notes; naively adding those notes fails because copy-forward text, templating, and repeated documentation swamp the genuinely new information at each timestep. The proposed framework strips temporal note redundancy before policy learning, comparing an embedding-space decomposition using singular value decomposition over local history subspaces against an interpretable sentence-level diff that filters previously documented sentences prior to encoding. On real intensive care unit data, the redundancy-stripped state representations beat both structured-only and raw-note baselines across all four off-policy evaluation methods tested (model-based rollouts, fitted Q-evaluation, weighted importance sampling, and weighted doubly robust evaluation).

Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL

Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Zhichao Shi, Hao Zhou et al. Training terminal-using agents with reinforcement learning requires executable environments whose rewards are trustworthy and whose difficulty sits where the current policy actually learns, but recipes like Self-Instruct and Evol-Instruct apply one prompting strategy to every seed regardless of policy state. Envs-FORGE turns verifier rewards into per-seed synthesis decisions: it estimates each seed's pass rate, scores six directions relative to a target learning frontier, and solves a small mixed-integer linear program to pick the action that conditions generation, which then rewrites the instruction, fixtures, oracle solution, tests, and Docker image together so only gold-verified bundles enter training. On Qwen 3.5 35B it lifts Pass@1 by 9.2 points on tb-core and 6.4 on tb-2.0, beating the best fixed recipe by roughly two points, and reaches 77.1% on SWE-bench Verified against 73.4% for the base model. All compared methods export 100 verified environments at comparable token cost, holding training-set size fixed.

CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving

Anisa Saleem, Duksu Kim cross-listed Long-horizon goal-directed driving asks a reinforcement learning policy to learn several competing behaviours simultaneously — reaching a distant goal, following a route, avoiding obstacles, obeying traffic signals — and a fixed reward gives no ordering over them. CORAL advances two schedules in lockstep: a five-stage curriculum that lengthens routes and tightens behavioural constraints, and a stage-aware reward whose weights shift from mission progress toward route following, safety, smoothness, and rule compliance as difficulty rises. The policy is a multi-stream actor-critic trained with Proximal Policy Optimization in CARLA on a compact 99-dimensional state — a polar LiDAR histogram plus telemetry, route geometry, and rule indicators, with no point-cloud encoder or bird's-eye-view raster. It succeeds in all twenty evaluation episodes on the hardest routes where two PPO baselines reach 5% and 10%, drops to 55% with both schedules disabled, and transfers zero-shot from its training town to seven unseen towns at 68–98% success.

Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View

Yixian Xu, Yuanrui Zhang, Shengjie Luo, Liwei Wang, Di He Reinforcement learning post-training for diffusion models has split into two apparently unrelated camps — reverse-trajectory methods based on discretized likelihood ratios, and forward-matching methods trained on reward-labeled noised versions of rollout samples. Starting from the regularized diffusion-RL objective and applying importance sampling between sampling stochastic differential equations yields an explicit policy-gradient estimator on trajectory space that contains the Itô integral behind Flow-GRPO-style updates, and an equivalent variance-reduced value-gradient form that reproduces the forward-matching structure of AWM and DiffusionNFT. The empirical gap between the two families is therefore a variance-reduction effect, not a difference in underlying RL principle. The resulting design space, organized by value-gradient estimation, weight function, and sampling choice, yields a multi-sample kernel density estimation value-gradient estimator with scale-bounded weights that improves on prior baselines on SD3.5-M and Qwen-Image.

Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training

Hanfeng Lu, Tianyu Feng, Suyi Li, Yuheng Zhao, Wei Gao, Shaopan Xiong et al. On-policy reinforcement learning runtimes execute rollout, reference scoring, and actor training as strictly serial phases, which wastes capacity for vision-language models where processing dense video inputs and prompt prefixes consumes much of each phase. Rollplex observes that prefix computation does not depend on the generated response, so it can overlap with rollout decoding without breaking synchronous on-policy semantics; making that schedule fit memory requires phase-aware control of high-bandwidth memory residency plus parallelism-aware weight sharing that reuses physical storage for tensors whose layouts are compatible across different tensor-parallel degrees, rebuilding only the incompatible ones. Colocating Qwen2.5-VL-32B naively would need roughly 165 GiB per GPU; on 32 H800 GPUs the runtime instead delivers 1.23×–1.30× speedup over serial colocation and 1.57×–2.24× over disaggregation at the same GPU budget.
3 more specialized papers

Safety & Alignment 16

Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation

Zhe Liu Large Audio Language Models (LALMs) are increasingly used for speech recognition and audio question answering, but checking whether they serve demographic subgroups equally is confounded by what is actually said and by speaker-specific vocal traits. The proposed evaluation framework is a semantic-aware mixed-effects regression that adds sentence-level semantic embeddings of the reference text as covariates and treats speaker identity as a random effect, with the embeddings drawn from the same LALM being evaluated so that semantic variation is controlled as that model perceives it. Across simulated data and real-world benchmarks, the method substantially reduces spurious fairness findings and produces subgroup performance gaps that are more stable and easier to interpret.

Language-Specific Gaps in AI Safety Training Datasets

Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, Rashidat Sikiru cross-listed Model providers cite multilingual safety benchmarks covering a dozen or more languages as evidence of safety for non-English speakers, but collection-level coverage claims can hide weakness in any individual language. An audit of 21 resources across 25 language slices, spanning Hausa (low-resource), Swahili (mid) and French (high), inspects provenance, annotation reliability, access, harm-taxonomy coverage and data reuse one language at a time. Gaps only partly track resource tier: within a single pipeline the Hausa slice fell below its own paper's translation-quality acceptance threshold while the Swahili output cleared it comfortably, and self-harm and sexual-content categories had no native-language coverage in either African-language tier. The authors argue this thinness lines up with the known persistence of multi-turn jailbreaks in non-English languages, and release a reusable slice-level audit protocol plus the safety-slice-audit dataset.

Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation

Ajay Pravin Mahale (Hochschule Trier) The EU AI Act obliges providers of high-risk systems to document how their systems reach decisions, and circuit discovery in mechanistic interpretability is the natural source of such evidence — provided two competent analysts using the same tool with different defensible settings would file the same thing. A pre-registered grid crossed seven analytic axes, each level drawn from a published implementation, over GPT-2 small on the indirect object identification task, mapping every discovered circuit through a deterministic claim map into a structured Annex IV statement. Across 15,840 specifications, 7,561 of which produced a claim, the derived statement flips across 73.2% of specification pairs and the most common claim covers only 41.1% of the space; standardizing the evaluation metric leaves the flip rate at 59.4%, and dropping circuit size from the claim entirely still leaves 27.1%. The underlying circuits are near-disjoint (median pairwise Jaccard overlap 4%) and functionally uncorrelated (Cohen's kappa 0.015), so this is not one mechanism described in different words; the study covers one model and one task.

ASSERT: A Measurement Pipeline for GenAI Audits

Riccardo Fogliato, Abhinav Palia, Xiawei Wang, Emily Sheng, Chad Atalla, Jean Garcia-Gathright et al. Audits of generative AI systems typically report a single compliance rate that gets used to compare systems, track regressions, and gate deployment, even though that rate reflects the auditor's measurement choices as much as the system itself. ASSERT is a specification-driven pipeline that helps draft a behavioral rubric and test cases, runs the audit, and binds every reported rate to a written record of the choices that produced it. In a case study on conversational deception, varying the dialogue setup, the simulated user, the judge, or the evidence bar for non-compliance shifts the reported rate enough to reorder which systems look better, which is exactly what the attached specification makes attributable.

Mandato: Protocol-Level Enforcement of Digitally Signed Mandates on AI Agent Actions with Cryptographically Chained Audit Trails

Giovanni Racioppi Agents now act on real systems through tool-calling protocols such as the Model Context Protocol (MCP), yet whether a given call was actually authorized lives in application code that is unsigned, unauditable, and logged without evidentiary weight. Mandato is a transparent MCP proxy that enforces digitally signed mandates — machine-readable artifacts stating which tools an agent may call, under what parameter and context constraints, for how long, and on whose behalf — evaluating each call against the mandate chain, blocking non-conforming ones inline, and writing every permit and deny decision into an append-only hash-chained audit log anchored with qualified timestamps. The mandate model is deliberately shaped after the civil-law delegation of authority so lawyers and auditors can read it, and the work maps the mechanism onto EU AI Act Articles 12 and 14, GDPR accountability, NIS2, and eIDAS 2; the reference implementation is described alongside a planned quantitative evaluation of overhead and tamper-evidence cost rather than measured results.

Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers

Thiago Sandoval, Ufuk Topcu Safety classifiers shipped alongside large language models tend to enforce the policy they were trained on rather than the one a deployer wants, and they decay as traffic shifts. Regime-Conditional Verification (RCV) wraps an off-the-shelf classifier without retraining it: from the classifier's internal representations it estimates the probability that each prediction disagrees with the deployer's policy, selectively corrects the likely-wrong ones, and reuses the same estimates as a label-free distribution-shift detector that triggers fine-tuning only when cheaper repairs fail. Across three classifiers and two benchmark datasets it improved policy adherence in every combination, catching up to 0.81 of previously missed unsafe content with the underlying classifier untouched, and it flagged all ten held-out attack campaigns in a deployment study.

BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs

Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos, Karin Sevegnani, Emine Yilmaz Work on social bias in language models mostly inspects final answers, leaving open which internal reasoning steps actually generate a biased outcome. BiasTrace is an annotation scheme that labels behaviours inside model-generated reasoning traces — both bias-specific ones such as unsupported demographic assumptions and general patterns such as overthinking — and links them to biased outputs, scaled up with validated LLM-as-a-judge labelling to build a large annotated dataset. The analysis finds that biased outputs typically arise from subtle reasoning behaviours rather than explicitly biased language, that reasoning-level labels improve bias detection, and that the annotated behaviours can be targeted for inference-time mitigation.

Training Fair Tabular Foundation Models

Patrik Kenfack, Jesse C. Cresswell, Anthony L. Caterini, Samira Ebrahimi Kahou, Ulrich A\"ivodji Tabular foundation models (TFMs) now lead on tabular prediction via in-context learning and are being used for high-stakes decisions, yet their fairness behaviour is largely uncharacterized. FairTFM builds fairness into TFM training itself so predictions come out fair in a single forward pass, working around two obstacles: sensitive attributes are rarely available in training data, and standard fairness techniques assume task-specific training rather than in-context learning. The recipe combines synthetic fairness tasks with a gradient reversal layer that pushes the model toward representations invariant to sensitive attributes, and across 132 fairness tasks it improves fairness consistently while keeping accuracy competitive.

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

Anna Borisiuk, Andrey Savchenko, Alexander Panchenko, Elena Tutubalina Facts seen frequently during pretraining are memorized more deeply and resist removal longer, yet unlearning methods apply the same gradient pressure to every target regardless of how common it was in training data. AdaPop combines local token confidence with a per-fact exponent scaled by an external popularity proxy such as Wikidata sitelinks or an LLM judge, and uses a dual-ascent controller to retune the retain penalty each epoch instead of hand-tuning the forget-retain tradeoff. Across three model families and two benchmarks it leaks roughly 5x less forgotten content under paraphrased queries and 1.6x less under adversarial reformulations than competing methods, with hidden-state analysis showing forget-set representations move further from the original model while retain-set representations stay in place.

Detecting Contaminated Code-Generation Prompt Batches via Influence Functions

Francesco Quinzan, Noor Munir, Yishun Lu, Stephen Roberts Prompts can steer code-generating language models toward insecure implementations, and defenses built around known vulnerability patterns or a fixed threat model miss attacks they were not designed for. CodeSIFT avoids specifying vulnerabilities at all: it uses influence functions to measure the parameter-space influence of code a model generates, then applies a statistical test to decide whether a candidate batch of prompts deviates from a benign reference distribution. Across three open-weight code models from 3B to 7B parameters and two new benchmark datasets covering varied vulnerability types, it reached AUROC up to 0.98 at moderate-to-high injection rates with well-calibrated false positive rates, substantially beating static analysis baselines.

A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation

Dipankar Sarkar cross-listed Regulatory standards phrased as principles — "fair, clear, and not misleading", "deliver good outcomes" — resist encoding as binary rules, so language models are increasingly used as the judge, yet they are rarely tested for anything beyond accuracy. Principle-Bench supplies 168 cryptoasset financial-promotion scenarios mapped to two UK Financial Conduct Authority principles, with paraphrase, adversarial keyword-stuffing, and boundary perturbations written to a pre-registered rubric, and evaluates four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. Alongside it, Ceca (Calibrated Exemplar-Cluster Assessment) provides a calibrated assessor with per-exemplar counterfactual attributions. No method wins on all four axes, and a 120B model judge, the strongest on benign inputs, drops from 0.74 to 0.27 accuracy on keyword-stuffed Consumer Duty prompts — behaviour the authors call "compliance theatre" — while a judge from another model family agrees with it at only Cohen's kappa 0.16 on that split, localising the failure to the model rather than the corpus.

Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

Wei Zhao, Zhe Li, Peixin Zhang, Jun Sun Neuron-level jailbreak defenses promise surgical precision but in practice either suppress toxic semantics spread across many pathways (requiring a large intervention footprint) or rely on external classifiers that also damage utility-critical neurons, and both stay switched on for benign traffic. Tripwire instead runs per-neuron hypothesis tests with false-discovery-rate control plus a utility-specificity filter to isolate genuinely safety-specific neurons, then clamps them to their harmful-conditional mean activations to inject an internal "this input is harmful" signal that fires the model's own aligned refusal. The clamp ships as two provably equivalent modes: a detector-gated inference-time intervention and an offline bias-patch weight edit. Across four safety-aligned models and four attacks it is training-free and cuts average attack success rate to at most 2.0% while costing only 0.5% to 5.3% on MT-Bench.

Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice

Syeda Anshrah Gillani, Mirza Samad Ahmed Baig cross-listed As patients start asking chat assistants which doctor to see, those assistants become infomediaries that silently decide which physicians are visible, so the authors run a prespecified randomized audit of what actually drives the choice. Seven models (six open-weight plus gpt-4o-mini) picked among five synthetic family-medicine physician cards with independently randomized attributes across 3,024 choice sets, three patient personas, and nine paraphrases, producing 40,068 scored responses, with gender and ethnicity signaled through names as in correspondence audits. Reputation dominates — a rating rise from 3.9 to 4.7 adds 31.4 percentage points of choice probability and a fee rise from $90 to $190 subtracts 20.0 — but demographic parity still fails, with female-signaled and Hispanic-, South-Asian- and Black-signaled names gaining 1.3 to 2.9 points over White-signaled ones, and mere first-listed position worth about $11 in fee-equivalent terms. The models cited gender or ethnicity in at most 0.03% of their stated reasons, so transparency regimes built on model self-report would miss these tilts entirely.

Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers

Taenyun Kim, Edyta Bogucka, Daniele Quercia Moral preference elicitation — polling people on hypothetical dilemmas and training a policy on the aggregated votes — is often treated as a neutral way to align AI with public values, but three upstream developer choices shape the result: which features go to a vote, which voters are sampled, and how the question is worded. Across two phases with 809 participants and three deployment contexts (AI kidney allocation, AI agents standing in for absent workers, and generative depictions of the deceased), the authors measure how each stage moves the outcome. Morally relevant features did not transfer across contexts, preferences split by political ideology for roughly a third of features with some differences reversing sign, and question framing alone widened or narrowed ideological gaps by up to a full scale point. The conclusion is that voting-based alignment cannot deliver fairness by aggregation alone, and that each pipeline stage should be audited and disclosed.
2 more specialized papers

Vision 16

TRUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection

Sebastian Doerrich, Andreas Franz Schwab, Francesco Di Salvo, Shyam Nandan Rai, Hanh Huyen My Nguyen, Christian Ledig cross-listed Computer-aided detection systems for colonoscopy are usually trained and scored on curated lesion-centric clips, which omit the long stretches of healthy tissue and the procedural artifacts that dominate a real examination. TRUE-Colon is a benchmarking protocol that measures deployment-relevant behavior alongside localization accuracy, applied to Faster R-CNN, YOLOv8, YOLOv11 and RT-DETR across the curated SUN and PICCOLO sets and 60 unedited full-length procedures from REAL-Colon. Models trained only on curated clips collapse on full procedures, while procedure-trained models reject non-polyp content far better and still retain accuracy on the curated benchmarks — an asymmetry consistent across architectures. The Transformer detector gave the strongest sensitivity and earliest, most persistent detections, with the convolutional detectors competitive at higher throughput.

SDO: Subspace Deconflicting Operator for Multi-Adapter Composition

Zhongsheng Wang, Zhedong Lin, Qian Liu, Xinyu Zhang, Jiamou Liu Loading several independently trained adapters into one diffusion backbone is an appealing route to multi-character generation, but deploying them together causes identity mixing, attribute leakage between characters, and unstable scenes. Framing the interference in parameter space, SDO reconstructs each adapter's layer-wise low-rank update, extracts compact subspace signatures, scores pairwise conflict by output-subspace overlap, and applies a permutation-equivariant transformation that suppresses harmful shared directions while preserving identity-specific ones, mapping results back into ordinary adapter weights usable in existing inference pipelines. Experiments report improved identity fidelity and compositional stability, with the gap over naive joint deployment widening as more adapters are composed.

Post-training Quantization for Hybrid Iterative Generative Models

Jing Gao, Junyi Wu, Wei Wang, Yan Yan, Yao Zhao Hybrid generative image models that couple autoregressive and diffusion paradigms produce high fidelity but pay for it in iterative inference cost, and applying standard post-training quantization to them causes outright model collapse. Diagnosing the failures identifies two causes: excessive activation outliers that force an unwinnable trade-off between covering the outliers and preserving normal-range precision, and amplified anomalies where small quantization errors compound into a calibration-inference mismatch. HyGenQ counters these with hierarchical cluster decoupling, which isolates outlier channels through multi-stage clustering, and scaling recalibration, which rescales anomalies past the Gaussian bound instead of truncating them, successfully quantizing representative hybrid models to 8-bit weights and activations while outperforming existing baselines across model families.

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang, Jialin Cui, Qingshan Li et al. cross-listed Interactive video world models need to generate frames causally with very low latency while still responding correctly to keyboard and mouse input, which is hard to preserve when distilling a bidirectional generator into a one- or few-step sampler. ForgeWM is a four-stage recipe — domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching against a bidirectional teacher — producing students specialized for denoising budgets of 1, 2, and 4 steps, plus a dual-path deployment mode where the one-step draft is re-noised and refined at replay time. On paired Minecraft trajectories the students lead the compared systems on imaging quality, motion-profile agreement, action-sign accuracy, and mouse-control accuracy while reaching the lowest LPIPS, and replay-time refinement matches four-step quality while staying about three times closer to the trajectory the player actually experienced than regenerating from noise. The same recipe transfers to gamepad-controlled first-person shooter gameplay.

Adversarial Learning of Classifier-Free Guidance Schedules

Ashwini Pokle, Alexandre Galashov, Arnaud Doucet, Mauricio Delbracio, Valentin De Bortoli Text-to-image diffusion models use classifier-free guidance, or CFG, with a single fixed scale applied across every timestep, sample, and prompt, which is rarely optimal and can produce artifacts. The guidance scale is instead learned as a function of diffusion time, conditioning, and the current noisy sample by casting the problem as density-ratio estimation: a discriminator estimates the time-dependent log-density ratio between the true and guided marginals, while a small generator network predicts the state-dependent scale. The learned schedules beat both hand-designed CFG heuristics and prior dynamic-guidance methods on standard text-to-image benchmarks.

AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations

Ying Huang, Wencan Zhang, Brian Y. Lim cross-listed Metrics used to judge face editing and privacy-protection systems are meant to stand in for human perception, but current representation-based ones model behavior without cognitive structure and assume a single universal observer, so they misrepresent how different populations actually perceive facial similarity. AlignFace builds three findings from the cognitive psychology of face perception — dependence on featural and configural attributes, nonlinear psychophysical scaling, and own-group bias — directly into its architecture, pairing a vision-language encoder and gated cross-attention with a concept bottleneck over interpretable face attributes and a neural generalized additive model for their nonlinear contributions, and is trained on a new FACETS dataset. Experiments report significantly better alignment with the perceptions of human subpopulations than baseline and recent learned perceptual metrics, while keeping the reasoning inspectable.

QuaSAR: Quantization Compensation via Stable Activation-Aware Rank Truncation

Lin-Fa Lee, Yi-Yu Chang, Kuo-Hei Yeh Training-free post-training quantization methods often use a goodness-of-fit score to decide which layers get closed-form residual compensation, but at 4-bit weights and activations (W4A4) that gate conflates genuinely unpredictable quantization error with numerical breakdown of the solver. Rank-deficient input activations produce ill-conditioned or singular Gram matrices, yielding spuriously negative fit scores so that layers which could in fact be compensated are thrown away. QuaSAR replaces the solver with a parameter-free truncated pseudoinverse that drops collapsed directions before inversion, reaching 81.42% top-1 accuracy on ViT-B under W4A4 and beating both prior post-training and fine-tuning-based baselines; combined with joint low-rank compression it reaches 80.26% at 54.7 MB.

Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation

Nikolai R\"ohrich, Isabell Hans, Felix Krause, Bj\"orn Ommer cross-listed Text-to-image diffusion models offer no dial for continuous concept-specific control such as aesthetic quality, and they remain unreliable on structures needing local coherence like rendered text or hands. Using a new notion of concept-wise mutual information, the authors find that generation of specific structures is localized to distinct layers, then build Concept Guidance (CoG): quantify each layer's concept-specific impact, and steer denoising with a weighted combination of predictions produced while concept-relevant layers are skipped. The method needs no training, gradients, external models, or prompt engineering and improves several targets on PixArt-alpha, SD3, SD3.5, and FLUX.1-dev out of the box.

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

Shuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen, Zihan Mei, Manting Li et al. cross-listed Modern video generators can produce convincing footage of wars, disasters, and other emergencies, but existing detection benchmarks say little about how detectors behave on this kind of content or after it spreads online. RA-Bench pairs 1,830 real videos across 10 social-risk categories with 16,056 clips generated from four open-source and five closed-source generators, then evaluates seven traditional detectors, ten zero-shot multimodal models, and two multimodal large language models fine-tuned for the task. None of the three detector families generalizes consistently across the benchmark, generation quality and conditioning affect each family differently while source-level patterns stay stable across sampling seeds, and the videos that fool human judges are also the hardest for detectors — with social dissemination (re-encoding and sharing) making detection harder still.

Style or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision Embeddings

Rory Ashton cross-listed Frozen image embeddings from encoders like CLIP report high accuracy on classifying paintings by art-historical movement, but the usual random train/test splits place works by the same painter on both sides, letting a classifier win by recognising the artist instead of the style. Re-running the task under an artist-disjoint protocol — holding out every artist in turn across a balanced set of 320 paintings from four twentieth-century movements — drops 5-nearest-neighbour accuracy from 0.87 to 0.77, with the loss concentrated almost entirely in Surrealism (down twenty points) while Impressionism and Cubism hold steady. The pattern repeats across four encoders including a vision-only self-supervised model, which locates the effect in visual structure rather than language supervision.

Marionette: Predicting World States, Rendering Geometry, Painting Appearance

Zian Meng, Zhen Li, Chuanhao Li, Qiang Li, Kaipeng Zhang cross-listed Interactive game world models usually autoregress pixels or latents directly, which forces pose, geometry, and occlusion to be tracked implicitly and lets errors compound over long horizons. Marionette splits the job three ways: a two-stage autoregressive dynamics model predicts an explicit 276-dimensional 3D world state of articulated skeletons, metric root trajectories, and rotations; a zero-parameter graphics bridge turns that state into pose-control videos with world-space geometry and occlusion computed in closed form; and a control-conditioned video diffusion model paints photorealistic frames. Because behavior lives in the explicit state, it can be corrected there — left unconstrained, two generated characters drifted 21.2 m apart against roughly 5 m in recordings with a third of frames showing ground penetration, while adding just a terrain collider and a separation cap to the state cut penetration by 66% with no change to the observation model. Forcing a mismatched action stream shifted root-aligned joint error by 31% across 48 held-out segments, and routing appearance through the predicted state cost little visual fidelity (FVD 831 versus 799 for recorded poses).
5 more specialized papers

Multimodal 13

VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation

Jiarui Hai, Karan Thakkar, Ke Chen, Yunyun Wang, Jiaqi Su, Rithesh Kumar et al. cross-listed Systems that synthesize speech from a written description of a voice currently cover a narrow range of speakers and offer little control after generation, such as cloning a voice or adjusting its emotion. VoiceDesigner addresses both with a hybrid data pipeline that combines digital signal processing with speech generation models to build a dataset spanning real and fictional voices, plus a diffusion transformer modified to handle complex conditioning across generation and editing in one model. Subjective and objective evaluations show better alignment with both voice descriptions and editing instructions than state-of-the-art text-to-voice systems, at comparable perceptual quality and usability.

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor, Shehzeen Hussain, Viacheslav Klimkov, Valentin Mendelev et al. cross-listed Speech language models mostly operate turn by turn and cannot handle a user talking over them, while duplex speech-to-speech systems that fix latency tend to sacrifice audio quality because recognition, interruption handling, and synthesis must all be optimized jointly. VoiceChat-TTS keeps synthesis modular: it consumes an LLM's text-token stream directly, emits silence when no text is available, and accepts explicit control tokens for interruption. The design gives always-on streaming speech at low latency and handles mid-utterance barge-in without resetting the KV cache, so the model does not restart generation when interrupted.

Content Based Video Narration of Gameplay with Vision Language Models

Mathew Varghese cross-listed Live esports-style commentary exists for professional broadcasts and almost nowhere else, so this system produces spoken narration for arbitrary gameplay recordings using a general-purpose vision-language model and text-to-speech, with no engine telemetry, game-specific instrumentation, or task-specific training. Three mechanisms carry it: temporal mosaic packing arranges nine sampled frames into one 3x3 image so an image-native model can reason about motion from a single payload, context-conditioned prompting replays the most recent narrations as assistant history to suppress repetition, and duration-conditioned generation with elastic alignment time-scales or pads synthesized audio to fill each segment exactly without a forced aligner. The speech stage can run fully locally via a 6-bit quantized 4B-parameter model on Apple silicon, the mosaic cuts per-minute image payloads by 9x, and the release includes a candid account of failure modes such as hallucinated game state, mosaic resolution loss, and prosody artifacts.

HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation

Yin Li, Ziyang Hu, Zhiyu Guo, Xiangyu Liu, Wenbin Li, Boo-Ho Yang et al. cross-listed Multimodal retrieval-augmented generation typically shreds structured documents into loose text and image chunks, discarding the section hierarchy and the local pairing between a figure and the text around it, which hurts both which evidence is picked and where images are placed in the answer. HAM-RAG carries document hierarchy through retrieval and generation as a grounding signal, keeping each piece of evidence tagged with its position and its neighboring text-image relations in the prompt, and ships HAM-Bench, a benchmark spanning game walkthroughs, web pages, scientific papers, and step-by-step recipes. Across several backbones it raises the main multimodal average by 17.3% over the strongest non-hierarchical baseline, with a 24.2% gain in image-text alignment on the game-walkthrough split.

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

Jinsheng Quan, Jianhua Li, Siyi Xie, Xuanke Shi, Kewang Deng, Zukai Chen et al. cross-listed Recovering geometry, matching points across views, and reasoning about spatial relations are usually handled by separate task-specific models or bolted-on geometry modules, which blocks any sharing between these complementary views of the same scene. SPARGen recasts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation inside one native multimodal generative model, emitting compact structured and linguistic answers as token sequences while producing dense geometric fields in image-aligned form so that spatial supervision shapes a single shared representation across all three task families. Evaluations across reconstruction, correspondence, and spatial reasoning benchmarks report competitive performance from that unified model.

Self-Supervised Visual On-Policy Distillation

Yijiang Li, Yijun Liang, Yunjie Tian, Bingyang Wang, Ke Zhang, Zhenfei Yin et al. cross-listed On-policy distillation for visual tasks normally needs an asymmetry between teacher and student — a larger teacher, ground-truth answers, or annotated regions of interest — which is unavailable when no privileged signal exists. Self-Supervised Visual On-Policy Distillation (S$^2$VOPD) inverts the setup by removing information from the student instead of adding it to the teacher: the teacher's distribution conditioned on the original image is distilled on-policy into the student's distribution conditioned on a strongly augmented view of the same image. A sweep over augmentation families shows asymmetry is what matters (symmetric self-distillation hurts), strength peaks at moderate levels, and augmentations that erase the question-relevant evidence produce large but useless discrepancies. Across six fine-grained perception benchmarks it lifts Qwen3.5-4B from 70.7% to 77.4%, beating open-source models up to Qwen3-VL at 235B as well as GPT-5.4, and recovering 96% of the gain from privileged-information methods.

Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding

Jeongwan Shin, Jaehyeon Kim, Donguk Ko, Jaeho Choi Millimeter-wave (mmWave) radar sees through darkness and occlusion, but pairing it with language models has been blocked by scarce radar-text data, incompatible dataset conventions, and the absence of a foundational radar encoder. The approach sidesteps all three with a minimal textualization interface that serializes each mmWave point cloud into short natural language so off-the-shelf LLMs can answer questions about it, and packages this as mmWave-QA, the first benchmark for language-conditioned mmWave human perception, built by harmonizing heterogeneous public datasets through calibration-aware preprocessing and a shared taxonomy across six scenarios and five question-answering tasks. Evaluation of current LLMs on the benchmark indicates non-trivial zero-shot reasoning over radar data and robustness where visual sensing degrades.

Seeing Red, Thinking Bad: Color Bias in Vision Language Models

Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh cross-listed Vision language models are being used for screening and recommendation tasks where text often arrives as an image, raising the question of whether visual styling changes their reading of semantically identical content. Stealth Visual Prompts alter only the color and contrast of rendered words while preserving wording, and applying them systematically showed that coloring positive words green consistently pushes sentiment predictions positive, with models frequently failing to register negative words in the same text. Latent analysis ties the shift to color-induced changes in the vision encoder's representations, and lowering text-background contrast increases reliance on visually salient cues and produces more incorrect visual question answering outputs.
5 more specialized papers

Robotics 12

Active Perception for Embodied Disambiguation

Yiwei Liu, Luwei Yang When a robot is told to fetch something and cannot tell which object is meant, the missing information is often physical rather than linguistic — the target is occluded, the viewpoint is bad, or a label is unreadable — so asking the user cannot resolve it. The proposed framework makes moving to a better viewpoint the primary information-gathering action, with a vision-language model deciding from accumulated visual evidence and dialogue history whether to keep observing, request clarification, or commit to a target. Real-robot experiments show physical information gathering and user-intent clarification working within one disambiguation loop, with active observation also improving the quality of clarification questions by revealing object names and attributes.

hint$^2$: Hierarchical World Models for Inference-Time Temporal Logic Guidance

Moritz Zoellner, Anastasios Manganaris, Ahmed H. Qureshi, Rohan Paleja cross-listed Linear Temporal Logic (LTL) can express the temporal structure and safety constraints that language-conditioned manipulation policies handle poorly, but LTL is evaluated over long trajectories while modern policies emit short action chunks and replan in closed loop. hint2 bridges that gap at inference time with two world models at different abstraction levels: a high-level model predicts how candidate actions change task-relevant atomic propositions to drive progress through the LTL automaton, and a low-level dynamics model predicts immediate state evolution for local safety guidance. It outperforms existing LTL-guided diffusion and inference-time steering methods on CALVIN, satisfies instructions with combined liveness and safety constraints, and transfers to a real UR5e arm.

Coverage Aware Active Evaluation for Failure Discovery with Paired Systems

Anjali Parashar, Rachel Luo, Apoorva Sharma, Sushant Veer, Edward Schmerling, Carson Sobolewski et al. Autonomous systems fail rarely and in varied ways, so finding those failures under a limited real-world testing budget is hard, and cheap proxies such as simulators or lower-fidelity policies surface failures that often do not transfer. The proposed method learns a local predictor of target-system risk by correcting proxy failure signals with control-variate-inspired residual modeling, then selects scenarios using a support-aware mutual-information objective that favors realistic, well-supported regions while spreading coverage across distinct failure modes. On autonomous driving, manipulation and quadruped velocity-tracking tasks it discovers up to twice as many failures as random sampling and active-learning baselines, including severe modes the baselines miss entirely.

AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning

Zhiyue Zhao, Jingyi Wu, Hairuo Liu, Mingyu Liu, Liyang Li, Hengdi Zhang et al. cross-listed Dexterous manipulation policies are expensive to train because robot teleoperation data is scarce and action spaces differ across hands and grippers, and models trained on mixed sources tend to tie task cues to embodiment-specific appearance. AdvDex is a Vision-Language-Action framework built on three pieces: OmniShare, a large multimodal dataset of human manipulation demonstrations with kinematic and tactile supervision; a Joint-Aligned Action Space of an SE(3) wrist pose plus 15 finger joints that puts human hands, robot hands, and parallel grippers in one representation; and domain-adversarial training that strips embodiment cues from the visual encoder. Experiments on hand-action prediction and real hardware report gains over baselines, zero-shot transfer of skills learned from humans to robots, generalization to unseen objects and scenes, and few-shot adaptation from little data.

Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use

Yi Ding, Yanzhao Yu, Xili Dai, Xianbiao Qi, Peiwen Sun, Xueqian Wang et al. cross-listed End-to-end Vision-Language-Action models must solve for continuous low-level actions directly, which makes them data-hungry and brittle outside their training distribution. ART (Agentic Robot with Tool-use) is a tool-injection layer that fine-tunes any such model to call off-the-shelf modules for low-level vision, high-level affordance estimation, and embodiment-specific control, shrinking the action space the policy itself has to search. Trained on 30,000 tool-use trajectories — far fewer than baselines consume — with a regimen aimed at long-horizon tool reasoning, it reports a 20% higher success rate than mainstream baselines in simulation and on real tasks such as pick-and-place in the dark from novel viewpoints, while allowing lighter deployment and incremental addition of new tools.

AgilePE: Autonomous UAV Pursuit-Evasion via Self-Play Reinforcement Learning

Wenhao Tang, Tianyang Chen, Zhejun Cui, Boyuan An, Jiayu Chen, Ruize Zhang et al. cross-listed Aerial pursuit-evasion demands fast decisions under coupled flight dynamics against an opponent whose behavior keeps changing, which rule-based and differential-game methods handle poorly at high dimensionality. AgilePE trains a policy that maps onboard state observations straight to collective thrust and body rate (CTBR) commands — no trajectory planner or waypoint controller in between — using competitive self-play with Prioritized Fictitious Self-Play (PFSP) against a diversified pool of historical opponents to stabilize optimization. A simulation pipeline modeling actuator response, communication latency, and domain randomization lets the learned policies transfer zero-shot to real quadrotors without task-specific tuning, with hardware runs reproducing the rapid dodging and flanking tactics seen in simulation, including two agents deployed against each other.

Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation

Yuxuan Chen, Wanruo Zhang, Xiao Li cross-listed Vision-Language-Action (VLA) models for robot manipulation are almost always benchmarked on static tasks, leaving open how they behave when the world moves and inference latency matters. ReflexBench supplies six dynamic, reaction-critical tasks along with an evaluation harness that decouples simulator stepping from robot control so that configurable latency can be imposed under both synchronous and asynchronous inference. The accompanying ReflexVLA model adds latent future prediction and multi-frame temporal fusion inside the vision backbone and cuts deployment latency with batched visual encoding and CUDA Graph replay, all without large-scale robot-data pretraining. It improves dynamic manipulation while staying competitive on standard static benchmarks, with real-robot experiments under practical deployment conditions.

Expected Free Energy-based Informative Path Planning for Robotic Mars Exploration

Ajith Anil Meera, Pablo Lanillos, Wouter Kouw cross-listed A robot surveying unknown terrain — hunting for water on Mars, say — must simultaneously map the environment accurately and reach the highest-value regions fast, all while paying for travel distance and each measurement; classical information-seeking and reward-seeking objectives each optimize only one side. The proposal adopts Expected Free Energy from active inference as a single action-selection criterion, maintaining a Gaussian-process belief over the information field and planning continuous trajectories that minimize expected free energy subject to hard path-length budgets. Across multiple simulated realizations this produces accurate posterior maps and finds top-value regions at the same time, beating information-theoretic baselines under matched settings, with few parameters to tune.
4 more specialized papers

Reasoning 4

Think in Latent, Explain in Language: Self-Explainable Latent Reasoning

Dayuan Zhao, Shengcao Cao, Yu-Xiong Wang, Liang-Yan Gui Latent reasoning compresses chain-of-thought (CoT) into compact embeddings for efficiency, but the resulting thinking is opaque: methods like Coconut act as black boxes, while Heima bolts on a separate post-hoc decoder that detaches the explanation from the computation actually performed. Self-Explainable Latent Reasoning (SELR) trains one model against two objectives at once — an answer loss shaping the latent trajectory toward correct answers, and a CoT loss teaching the same model to decode its own latent representations back into human-readable reasoning steps. Validated on both large language models and vision-language models, SELR reports better token efficiency and accuracy than baselines while producing its own explanations with no auxiliary decoder.

Capacity-Dependent Effects of Data Selection for Reasoning

Cuong Dang, Hoang Anh Just, Ruoxi Jia Recent work on selecting supervised fine-tuning data for reasoning suggests picking teacher responses that the student already assigns high likelihood, on the theory that closer-to-distribution supervision teaches more efficiently. Controlled experiments on mathematical reasoning with students from 1.5B to 8B parameters show that this preference is not universal but capacity-dependent: high-likelihood data gives faster, more stable early gains especially for small models, while low-likelihood data becomes increasingly valuable for larger models trained longer. Analysis of learning dynamics finds small models fail to absorb low-likelihood supervision and lapse into shallow or repetitive behavior, whereas larger ones move toward the teacher distribution under it, and a capacity-constrained account of distillation explains how data difficulty, data span and student capacity interact. The practical conclusion is that data selection should be conditioned on model size and compute budget rather than a fixed likelihood preference.

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig Reasoning-oriented training can make a model's traces look more deliberative without amplifying the behaviors that actually correlate with getting the answer right. Behavioral Lift measures how much correctness changes when a given behavior is present versus absent in a trace, applied to 15,282 annotated traces across 15 models and 6 text-only and vision-language benchmarks using a taxonomy defined for both kinds of trace. The result is an amplification-lift gap: thinking models heavily amplify self-correction, hypothesis testing and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment and self-awareness. Uncertainty acknowledgment is amplified 3-7x yet is weakly or negatively associated with correctness, and confidence calibration is among the strongest correctness signals in both modalities yet is barely amplified, arguing for process-level training objectives that reward calibrated, grounded reasoning rather than surface form.

MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement

Lushi Pu, Weiming Zhang, Xinheng Xie, Zixuan Fu, Bingxiang He, Hengyu Zhao et al. Autoformalization — turning natural-language mathematics into machine-checkable code such as Lean 4 — fails when models must recall the type and definition hierarchy of a large library like Mathlib purely from parametric memory, and typical data pipelines just filter single-pass generations with no revision step. MathForm adds a retrieval planner that gathers relevant Mathlib definitions and prior formalizations before generation, then revises outputs using compiler diagnostics and semantic-consistency feedback; this pipeline produced FormalVerse, roughly 367K verified Lean 4 statements. The resulting MathForm-8B, trained with supervised fine-tuning plus reinforcement learning, reaches average Pass@8 of 88.06% on syntax checks and 72.37% on consistency checks across six benchmarks, outperforming several specialized 32B autoformalizers at a quarter of the size, including 63% and 37% consistency pass rates on the hard FATE-H and FATE-X subsets.