Monday, August 31, 2026
Highlights
Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap
Post-training quantization is usually assumed to preserve behavior, so models are certified at full precision and compressed afterward without re-evaluation — a workflow this work formalizes as a validation-deployment gap using Quantization Behavioral Equivalence Classes, proving that class membership does not imply behavioral equivalence. A three-stage adversarial fine-tuning procedure embeds payloads that stay dormant under source-precision checks but activate under INT8 or 4-bit compression, demonstrated on multilingual encoder-decoder translation and political stance classification rather than just decoder-only models. Backdoored translation models go from zero measured friend-foe corruption at repaired FP16 to up to 85.02% inversion after quantization, and cross-quantizer analysis shows persistence depends on the specific quantization scheme and architecture rather than nominal bit-width, arguing that the deployed configuration itself must be behaviorally certified.
Post-training quantization is usually treated as a semantically neutral optimization, but because it is a many-to-one map from full-precision weights to discrete bins, a model can be engineered to pass full-precision safety checks and only turn malicious once compressed for edge deployment. The paper formalizes this as Quantization Behavioral Equivalence Classes — the set of FP weight vectors that collapse to the same quantized model — proves that class membership does not imply behavioral equivalence, and uses that structure to build backdoors triggered by LLM.int8() or NF4 compression alone.
- The three-stage pipeline fine-tunes on a corrupted dataset, computes per-weight QBEC intervals for the target quantizer, then runs projected gradient descent on clean data clipped inside those intervals, yielding a checkpoint that behaves benignly at FP16 while quantizing back to the malicious model.
- On English→Ukrainian tactical translation, repaired models show zero measured friend-foe corruption at FP16 but reach 82.27% inversion for
NLLB-200-1.3BunderNF4and 85.02% forM2M100-1.2B, with BLEU falling only about 5–6 points (22.87→17.17 and 33.42→31.30). - The same recipe applied to political summarization with the
PoliTunedataset produces a measured ideological shift of ΔBias up to 0.33 forLlama-3.2-1Band 0.24 forGemma-3-1B, while MMLU on political and legal subtasks drops only 4–7 points. - Cross-quantizer transfer is governed by quantizer geometry rather than bit-width: an
INT8-optimized attack onNLLBretains 55% of its strength underNF4(T=0.552) but collapses to T=0.061 underINT4andFP4, and excluding shared embeddings and the output head from repair liftsNLLB'sNF4persistence from 46.33% to 71.75% with no measured BLEU change — a warning that partial repair is not automatically the conservative choice. - The evidence is bounded: all numbers come from single fixed-seed runs with no confidence intervals, the translation corpora are synthetic and may exaggerate how learnable the lexical swaps are, models are all around 1B parameters, and the authors are explicit that BLEU and MMLU are stand-ins for a real audit rather than proof that stronger source-precision checks would miss the backdoor.
Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess
Searchless chess networks such as Leela Chess Zero's Chessformer reach master strength in a single forward pass by distilling Monte Carlo Tree Search visit counts, but imitating a search is a poor proxy for playing without one. Fine-tuning with self-play reinforcement learning, the usual entropy bonus (a reverse Kullback-Leibler divergence toward uniform) is swapped for a mass-covering forward KL toward the network's own MCTS prior, paired with a sampling temperature that sharpens once the value head is confident about the outcome. In roughly two thousand steps this lifts puzzle accuracy from 93.9% to 94.9% and mate-in-four from 77% to 81% without losing playing strength, while a control fine-tuned on puzzles alone posts the largest tactical gains yet loses about 260 Elo — a better puzzle-solver is not a stronger player. Without any regularizer, self-play collapses onto a single line of play.
Searchless chess networks are trained by distilling an AlphaZero-style search's visit counts, but imitating a tree search is a mismatched objective for a policy that must commit to one move from a single forward pass. The fix here is self-play reinforcement learning on top of Leela Chess Zero's released Chessformer, with the standard entropy bonus replaced by a forward, mass-covering KL divergence toward the network's own MCTS prior — "prior-directed exploration" — so the exploration budget covers the moves the prior already judges promising rather than every legal move.
- The recipe is deliberately minimal: a single-step policy gradient with a TD(0) value bootstrap from exponential-moving-average target weights, no replay buffer and one optimizer step per freshly collected batch (which makes PPO clipping inert), plus an entropy-adaptive sampling temperature that sets each move's temperature to the value head's win/draw/loss entropy so self-play sharpens once a position is decided but stays stochastic while the outcome is live.
- In about 2,000 steps accuracy on a
100,000-puzzle Lichess suite rises from 93.85% to 94.87%, mate-in-four from 77.2% to 81.4%, and searchless strength holds or slightly improves at +16.6 ± 8.3 Elo over the2374-rated base — a real but slender edge (361–318–321 over 1,000 games, likelihood of superiority 0.94). - Tactical accuracy and playing strength dissociate under matched compute: all self-play arms land within a one-point accuracy band while their ratings straddle the base, and a supervised control fine-tuned on puzzles alone posts the study's best tactics (96.80%) while shedding roughly 260 Elo, so a better puzzle-solver is not a stronger player.
- Distribution-level diagnostics explain what the anchor buys — unregularized self-play collapses onto a single opening line carrying 99.9% of its trajectory mass, the entropy bonus keeps the argmax right while draining probability off winning lines (geometric-mean solution probability 0.222 versus the base's 0.660), and the mass-covering forward KL retains the hardest solutions best (0.236 versus 0.199 for the base and 0.178 for a mode-seeking reverse-KL anchor on puzzles rated ≥2400).
- The caveats are stated plainly by the authors: the strength gain is modest and partly substitutable with the adaptive temperature alone (which already reaches +9.8 Elo on its own), the forward-KL arm is statistically tied with the reverse-KL anchor on rating, the headline is the best of many swept configurations with one run per cell so some margin is winner's-curse inflation, and the gains are retention-gated — fine-tuning promotes near misses the base ranked second or third but recovers nothing whose solution the base had already pushed below about
0.05probability.
Fast Weight Attention for Continual Learning
Recurrent fast-weight memories and selective state-space models squeeze an expanding context into a fixed-size state, which makes the state transition an online learning rule; under read-after-write autoregressive semantics the correct local training example at each step pairs the previous key with the current value, not the same-step key, and the common same-step association optimizes a different internal objective. From that alignment the authors derive normalized first-order updates for squared-error regression — Falcon-1 (a scalar normalized least-mean-squares rule), Falcon-2 (per-column) and Falcon-3 (sliding-window mini-batch) — plus inner-product variants, each with recurrent, masked-parallel and chunk-parallel forms and a numerically stable positive-decay renormalization. Representative variants stay competitive on language modeling and improve length extrapolation on variable-digit addition, with the framework separating temporal alignment, plasticity, forgetting and bounded rehearsal into independent knobs.
Recurrent fast-weight memories and selective state-space models compress an unbounded context into a fixed-size state, which turns the state transition into an online learning rule — but which training example that rule actually sees at each step has been left implicit. The core claim is that under read-after-write autoregressive semantics the correct prefix-aligned example at step t is (φ(k_{t-1}), v_t), not the conventional same-step pair (φ(k_t), v_t), and that fixing this alignment yields a principled family of update rules.
- The framework derives normalized first-order updates from two objectives — squared-error regression and a negative inner-product loss — giving
Falcon-1(a scalar NLMS step),Falcon-2(its per-column extension), andFalcon-3(a sliding-window mini-batch update), each with an inner-product counterpartFalcon-1A/Falcon-2A/Falcon-3A. - The same-step association is not wrong in the causal sense, it simply optimizes a different internal objective, so the shift to prefix alignment is a semantic correction rather than a fix for a leakage bug.
- Each variant is given recurrent, masked-parallel, and chunk-parallel formulations plus a numerically stable positive-decay renormalization, so the rules stay trainable at scale rather than being confined to sequential simulation.
- Empirically the reported gains are language modeling parity with existing recurrent baselines and improved length extrapolation on variable-digit addition, a task chosen because it isolates whether the memory generalizes past the training sequence lengths.
- The main limitation is evidential rather than structural: only representative variants are evaluated, the headline results are stated qualitatively without accompanying perplexity or accuracy figures, and addition is a narrow probe from which broader length-generalization claims do not automatically follow.
Rubric-to-Code Credit Assignment for Reinforcement Learning
Generating an interactive web application involves many distinct user-facing requirements, each tied to a localized region of code such as an event handler, state update, or CSS selector, but GRPO compresses all of that into one sequence-level reward applied uniformly across every token. Rubric-to-Code Credit Assignment (RCCA) builds training tasks around explicit functional rubrics, uses a hierarchical reward that separates format, source-code, runtime, and functional failures, and aligns evaluator-written textual attributions with the code spans and tokens responsible. The resulting Ling-RCCA-Flash scores 41.25 on MiniAppBench, a 32.20-point gain over the Ling-3.0-Flash base model and slightly ahead of Claude Opus 4.5, and reaches 76.19 on ArtifactsBench, topping the official leaderboard setting.
Interactive web app generation is judged by many user-facing behaviors at once, but standard GRPO collapses those outcomes into one sequence-level reward and spreads the advantage evenly over every token, so the model learns whether an app worked without learning which lines made it work. RCCA keeps the rubric structure intact: it grades applications through staged validity checks and then routes the evaluator's textual diagnostics back onto the specific code spans responsible for each pass or failure.
- Each training task pairs a natural-language request with a set of independently checkable rubrics, split into initial-state requirements verified after page load and dynamic requirements that require executing an interaction path, and the same rubric set drives task construction, reward computation, and failure attribution.
- A gated hierarchical reward scores format and source-code failures at 0, runtime failures at 0.1, and otherwise computes a rubric score in [0.2, 1.0] by subtracting penalties that scale with each violated rubric's importance tier (core functionality, behavioral correctness, or rendering quality) and severity, with score ceilings so a severe core failure can't be offset by satisfying minor requirements.
- Token-level credit assignment maps evaluator-identified code regions — event handlers, state updates, DOM fragments, CSS selectors, expanded along enclosing functions, event bindings, and state reads/writes — to completion tokens via tokenizer offset mapping, then weights the
GRPOloss asymmetrically: on negative-advantage samples, directly implicated tokens get weight 3.0 and contextual ones 1.5, while on positive-advantage samples those regions stay at 1.0 and unrelated code gets 1.2, so error regions are punished sharply without being reinforced on success. - Trained on
Ling-3.0-Flash(a 124B hybrid-linear MoE with 5.1B active parameters),Ling-RCCA-Flashreaches 41.25% average pass rate onMiniAppBenchversus 9.05% for the base model and 26.85% after SFT — a 14.40-point gain from RL alone — edging pastClaude-Opus-4.5at 41.14%, with the largest margins on hard tasks (37.60%) and Lifestyle (70.00%); onArtifactsBenchit scores 76.19, up 4.48 over the SFT checkpoint and 3.64 aboveGPT-5's official leaderboard entry. - The approach is confined to single-page HTML/CSS/JavaScript artifacts and untested on multi-page apps, backends, persistent storage, or auth flows, and because credit assignment rests entirely on evaluator-generated judgments and attributions, wrong diagnostics push gradient weight onto the wrong spans — most likely when a failure emerges from interactions among distant code regions rather than one local handler.
The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension
The question here is what geometry governs how well softmax attention can be approximated by a low-rank matrix, measured as the maximum-row-ℓ1 approximation rank that preserves every bounded output. Two sharp worst-case laws separate the effect of where keys and queries live: spherical self-attention has rank Θ(min{n, (1+β)^((d-1)/2)}), while full-ball geometry adds a radial degree of freedom and yields Θ(β^(d/2)) at large temperature. Within a fixed head, row-wise softmax cancels row-scalar logit directions, leaving a visible query–key interaction dimension r that gives a minimax-sharp per-instance exponent of r/2; a calibration set of 84 BERT-base heads shows modest effective-dimension reductions across many head and temperature settings, so support geometry sets worst-case temperature scaling while softmax-visible interaction geometry controls per-head complexity.
Softmax attention's rank complexity is usually discussed via head dimension or attention-matrix spectra, but row normalization erases entire logit directions and the geometry of the query/key support governs how sharply the induced Gibbs family can localize. The paper studies maximum-row-ℓ₁ approximation rank — exactly the least unrestricted rank that preserves every bounded vector-valued attention output to error ε — and separates two geometric roles: support shape sets worst-case temperature scaling, while the softmax-visible query–key interaction subspace sets per-head complexity.
- Two sharp worst-case laws pin the temperature exponent to support geometry: spherical self-attention has rank Θ(min{n, (1+β)^((d−1)/2)}), while full-ball geometry adds one radial degree and gives Θ(β^(d/2)) in the large-token regime
β ≥ β₀(d,ε),n ≥ C_d·e^(β/8)— so one radial degree of freedom shifts the exponent by exactly 1/2. - The upper bounds come from a weighted Gibbs-row cover whose constant is uniform in alphabet size and in all positive base weights, built by showing the log-partition function's one-homogeneous extension is a support function of a convex body sandwiched between
B₂^mand3B₂^m, then applying Dudley–Bronshteyn–Ivanov polytope approximation plus Pinsker; the lower bounds plant near-orthogonal or lattice-Gaussian "message" configurations and decode them into an identity block that forces rank. - For a fixed head, row-softmax quotients out row-scalar logit directions, leaving a visible interaction dimension
rwith the per-instance lawr_ε(A) ≤ min{n, C_r(1 + Γ/ε²)^(r/2)}whereΓ = βρ_Q ρ_K, and explicit bounded constructions in ℝ^(r+1) achiever_ε(A) ≳ β^(r/2), making the r/2 exponent minimax sharp. - A projective softmax perturbation bound (
‖softmax(x+h) − softmax(x)‖₁ ≤ 2 tanh(osc(h)/4)) converts discarded interaction into a uniform row errorτ_W, yielding a robust bound at residual budgetε − τ_Wand a reproducible tolerance-indexed SVD effective dimension whose criterion is output error, not explained variance. - Empirically, synthetic fits track the predicted slopes closely (sphere 0.504, 1.012, 1.381 vs 0.5, 1, 1.5; full-ball 0.476, 0.951, 1.427, 1.903 vs d/2), and on a fixed 84-head
BERT-basecalibration set with exact visible dimensions of 63–64, 41.9% of head–temperature cells admit a strictly smaller effective dimension at ε = 0.25, with Spearman associations of 0.574 and 0.606 against attention-SVD and representative-row rank certificates. - The main caveats are stated plainly: the
β^(d/2)law needs an extreme token count (the largest construction implies an effective token count around 10^88,946), the analytic constantC_rcan carry2^O(r)dependence and is numerically useless at r ≈ 50–64, the BERT study calibrates one model rather than establishing a universal learned-head law, and low approximation rank implies neither sparsity nor arithmetic speedup.
Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
Pretraining dense large language models on English-prevalent corpora, this study maps how optimal learning rate and batch size scale jointly and how each evolves marginally with model capacity and data budget, fitting a model of those relationships. Using a Warmup-Stable-Decay schedule, it also measures what learning-rate annealing buys across a wide sweep of settings and whether hyperparameters chosen for the stable phase transfer to the decay phase, then evaluates recently proposed loss scaling forms that explicitly model the capacity-data interaction. Those forms captured both undertraining and overtraining regimes well across the experiments, and the complete collection of pretraining runs is released open-source as a baseline for future OpenEuroLLM models.
Pretraining hyperparameter choices — learning rate and batch size — have to be picked before the expensive run starts, and most published scaling laws only describe their jointly optimal values rather than how each one moves independently as models and data grow. This study fits a model of both the joint and marginal scaling of learning rate and batch size for dense LLMs on English-prevalent corpora, then reexamines the loss-vs-capacity-vs-data relationship using scaling forms that explicitly model the interaction between the two.
- The experiments sweep learning rate and batch size across a grid of model sizes and data budgets under a
Warmup-Stable-Decayschedule, fitting not just the jointly optimal setting but the marginal evolution of each hyperparameter with model capacity and dataset size separately. - Because
WSDsplits training into a stable phase and an annealing phase, the authors quantify how much loss reduction annealing actually buys across a broad hyperparameter range, and test whether the optimum found during the stable phase still holds after decay — i.e. whether hyperparameters transfer between phases, which determines if cheap stable-phase sweeps are a valid proxy. - For the loss law itself, recently proposed forms with an explicit capacity-data interaction term are reported to fit both undertrained and overtrained regimes across the sweep, where classic separable Chinchilla-style forms tend to fit only one side well.
- The complete collection of pretraining runs is open-sourced, giving the
OpenEuroLLMeffort a reproducible baseline and scaling procedure rather than a set of borrowed constants. - The headline caveat is scope: the corpora are English-prevalent and the models dense, so transfer of these fits to the multilingual, potentially sparse models
OpenEuroLLMultimately targets is assumed rather than demonstrated, and the abstract reports no absolute loss or parameter-count numbers to anchor the fitted coefficients.
Sliding-window beats linear attention
Retrofitting pretrained LLMs with linear attention has been promoted as the fix for quadratic attention's ever-growing key-value cache, but the authors argue this line of work has never been compared against simple baselines. Testing sliding window attention (SWA) with attention sinks against post-trained linear-attention models across several LLMs and downstream tasks, they find SWA matches or beats them, and on long-context tasks (Needle-in-a-Haystack and BABILong) scores 2 to 10 times higher. Since SWA requires no post-training and is fast and memory-light, they recommend it over converting models to linear attention, which they suspect would need training from scratch or extensive post-training merely to match.
Linearizing a pretrained LLM to escape the quadratic KV cache is an active research line, but it has been benchmarked mostly against sink-free sliding-window attention, which is known to collapse catastrophically. The claim here is that the simplest baseline — swapping the attention mask for sliding-window attention with attention sinks, with no training whatsoever — matches or beats post-trained linear attention across 11 base models from 1.3B to 70B.
- The method is just an inference-time mask change,
SWA(w, s): each token attends to the previouswtokens (64 to 512) plus the first 4 "sink" tokens, requiring zero post-training tokens, noLoRAstage, and no specialized linear kernels sinceFlashAttentionalready supports it. - On short-context knowledge and reasoning,
SWA(64,4)recovers 93.2% of teacherMMLUand 99.0% of the 6-benchmark average with 0 training tokens, against 83.2%/97.5% forLoLCATsat 40M tokens and 92.4%/99.1% forQRWKV6at 350–700M tokens, and it wins the average in 9 of 11 model comparisons. - The gap widens sharply on long context with
Llama 3.1 8Bas base: at 4K onS-NIAH, SWA scores 17.2–23% whereLoLCATspeaks at 5.8% andLiger-GLAat 0.8%, and onBABILongat 4K SWA reaches 15% versusLoLCATsat 3%. - SWA is also the fastest decoder with throughput flat in context length, and its memory stops growing once the window fills — lower than linear attention at
w=64, though atw=512the KV cache exceeds a linear recurrent state. - The honest caveat is that every sub-quadratic option is badly degraded in absolute terms — full attention scores 100% on
S-NIAHat 4K and 60% onBABILongat 4K where SWA gets 23% and 15% — and the study covers only training-free SWA, no hybrid full-attention layers, and contexts no longer than 4K.
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
Long-horizon agentic tasks force LLMs to retrieve and integrate scattered information over many turns, so preserving full interaction history makes the working context grow without bound. Existing proactive context-management methods give models only search, deletion, and summarization tools, explore all editing actions as if their impacts were equal, and assign one trajectory-level reward to every intermediate edit. ContextPilot adds planning, long-term memory, and soft context-offloading tools, and trains with an RL method that uses context and entropy variation to identify critical editing decisions for branch sampling and then estimates action-level advantages from all branched trajectories passing through that edit. On long-context question answering and deep-search benchmarks it beats existing baselines across several base models while keeping a more compact working context.
Long-horizon agents that append every tool call and observation to their context hit a wall: the working context grows monotonically, and the usual fixes (rule-based truncation or summarization) give the model no say in what gets kept. ContextPilot from Tencent Youtu Lab and Shanghai AI Lab lets the agent manage its own context through an expanded tool set, and trains that behavior with a reinforcement learning recipe that treats context-editing decisions as first-class actions rather than incidental steps.
- The tool set extends the search/delete/summarize baseline of
StateLMwith planning tools, a structured long-term memory built from entities, timestamps and event episodes with edges between related items, and soft offloading tools includingsummarizeContext,compressContext(viallmlingua-2) and afoldHistoryoperation that discards history into keywords plus a summary recoverable by keyword search. - On the training side, context-aware partial rollout scores each context management action by how much it changed context length and token entropy relative to the initial query state, then spends the leftover rollout budget branching from the most sensitive actions — the paper collects 128 snapshots per query starting from 8 trajectory-level rollouts — and credit for an intermediate snapshot is the average reward over all terminal trajectories passing through it, an unbiased estimator with variance σ²/n instead of σ².
- Working within a 32K context window,
ContextPilot-14B-RLreaches a 72.20 average acrossNovelQA,∞Bench,LongMemEval-SandBrowseComp+, beating its own 128K-contextQwen3-14Bbackbone at 53.26, whileContextPilot-8B-RLgains 3.55 points overStateLM-8B-RL; on deep search the method adds 1.51 average points overSUPOacross theWebSailor-7BandWebExplorer-8Bbackbones. - Reinforcement learning helps most where contexts are longest — the gain over the supervised model is 5.34 points on
BrowseComp+(average input 552K tokens) versus a modest lift onNovelQA(119K tokens, and partly seen during supervised training) — and the ablation shows entropy-based branching alone is unstable, costing 1.32 points onBrowseComp+until context variation is added. - Per-turn input length stays flat at roughly 8K–10K tokens on
BrowseCompwhereWebExplorer-8Bgrows nearly linearly toward 30K, though the authors note the tool set still may not cover all context-editing needs, hyperparameters for partial rollout and credit assignment went unsearched for compute reasons, and evaluation never leaves long-context QA and deep search — agentic coding and GUI agents remain untested.
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
Factual question answering benchmarks assume a single canonical answer, which hides whether a model retains genuinely divergent accounts of long-tail facts. ElephantBench is a closed-book probe of 1,094 questions built by an auditable graph-based pipeline that retrieves related documents from a low-exposure web corpus, identifies naturally occurring disagreements, converts them into multi-account records, and verifies each answer against the source documents, authoritative public sources, and human annotators. Across 32 models, the strongest recovers both accounts on only 52.4% of questions, recalling one and omitting the other on nearly all the rest; larger models and inference-time reasoning improve recall without eliminating the incompleteness, and corpus analysis links completeness to how much exposure the minority account has.
Factual QA benchmarks that assume one canonical answer can't reveal whether a model's parametric memory holds the multiple competing accounts that real web sources report for long-tail facts. ElephantBench is a closed-book probe of 1,094 questions mined from naturally occurring cross-source disagreements, scoring not just whether a model recalls a rare fact but whether it recovers every verified account of it.
- The construction pipeline takes the documents a
DCLM fastTextquality filter discards, builds a document graph over them where support edges link agreeing sources and conflict edges mark incompatible values for the same subject–attribute pair, then turns each conflict-centered subgraph into a QA record — knowledge-point clustering plus entity matching cuts the candidate pair space by at least 93.4% versus exhaustive comparison, and every answer is checked by an LLM against source text, independently confirmed by a web agent on authoritative sources like Wikipedia, and finally audited by human reviewers. - Across 32 models, the best system
Kimi-K3recovers both accounts on only 52.38% of questions, withGemini-3.1-Proat 50.37% andGPT-5.5at 50.18%; the dominant failure is not amnesia but incompleteness, since failed recall for the top three models sits at just 2.19–2.65% while partial recall runs 45–47%. - Scale helps recall but not completeness —
Qwen3.5from 2B to 397B lifts complete recall from 1.65% to 32.27% and drops failure from 81.35% to 8.50%, yet partial recall climbs from 17.00% to 59.23% — and inference-time reasoning is similarly uneven, adding 13.99 points forGPT-5.6-Solbut costing smallQwen3.5models up to 0.64 points as deliberation converges on the consensus account and suppresses the minority one. - Corpus analysis ties the failure mode to exposure asymmetry: a one-standard-deviation increase in majority-account exposure raises partial recall by 14.18 points and cuts failure by 10.17, whereas the same increase on the minority side raises complete recall by 15.13 points — so overall fact frequency governs whether the model knows anything, while minority-side exposure governs whether it knows everything.
- The ceiling is shared rather than model-specific: an oracle pooling all 32 configurations reaches 81.2% complete recall and zero failures, but 206 questions remain partial for every system, and the authors caution that exposure results are associations from one public corpus rather than measurements of any evaluated model's undisclosed pretraining mixture.
Video Generative Models as Geometry Learner
Generative geometry estimation currently adapts pretrained image diffusion models, either training depth and normal predictors separately — forfeiting the correlation between the two targets — or jointly fine-tuning modified backbones, which needs a lot of labeled data. GeoNeXt instead repurposes a pretrained video generative model and casts geometry estimation as next-frames prediction, inheriting temporal structure and richer priors while adapting them to jointly model image and geometry targets in both directions. On zero-shot monocular depth and surface normal estimation across diverse datasets it beats prior task-specific and unified generative methods, and rivals discriminative state-of-the-art systems trained on over 100x more data, leading on several benchmarks.
Monocular geometry estimation with image diffusion priors forces a choice between separate per-task models and heavily modified joint architectures that need far more labeled data. GeoNeXt sidesteps both by repurposing a pretrained video diffusion model, casting depth and surface normal prediction as next-frames generation conditioned on the input RGB image.
- Building on
Stable Video Diffusion, the method encodes the image–depth–normal triplet through a frozenStable DiffusionVAE, replicates the image latent into the geometry slots as conditioning, and fine-tunes only the denoising U-Net with the text/CLIP branch removed — no new attention or switcher modules. - Trained on just 59K synthetic samples from
HypersimandVirtual KITTI 2, it beats the unified generative baselineGeoWizard(208K samples) on zero-shot depth by 6.2 AbsRel onKITTIand 10.4 AbsRel onDIODE, and posts the bestETH3DAbsRel of 5.6 against 13.1 forDepth Anything, which uses roughly 100× more training data. - On surface normals it is competitive with the specialist discriminative model
DSINEand the generativeE2E-FT, reaching mean angular errors of 16.4° oniBims-1, 33.0° onSintel, and 22.8° onOASISwhile predicting depth from the same forward pass. - Ablations isolate the two design choices that matter: co-generating the image alongside geometry rather than geometry alone, and joint rather than separate depth/normal training, each worth roughly 0.5–1.2 AbsRel and 0.5–0.9° mean angular error; swapping the depth/normal frame order changes results negligibly.
- Inference runs 5 denoising steps × 5-seed ensemble in 10 s at 768×768 on an A5000 — far cheaper than
GeoWizard's 272 s at 50×10 — but the model still predicts only affine-invariant depth, is trained purely on synthetic indoor and driving data, and the reportedGeoWizardnumbers are the authors' own reproductions.
Applications 80
Dandelion: A Spherical Flower for Neural Simulation of Planetary Dynamics
Scientific machine learning architectures are built for flat Euclidean grids, so applying them to spherical dynamics on a latitude-longitude grid distorts convolutions near the poles, makes 2D FFTs in Fourier neural operators falsely assume double periodicity, and warps geodesic distances in vision transformer positional encodings. Dandelion is a spherical version of the warp-based neural PDE solver Flower: each layer predicts a tangent-plane displacement and transports features along great circles, with hierarchical U-Net-style pooling done entirely in the spherical-harmonic domain, so spatial mixing comes from coordinate warps and not convolutions. A companion benchmark suite of natively spherical PDE datasets — a modified Galewsky jet, chained turbulence, Cahn-Hilliard decomposition, spherical Riemann shocks, Held-Suarez transport, and global ocean dynamics — fills the gap between toy problems and ERA5, and Dandelion places first or second on every dataset, with its margin over non-warp baselines growing at higher resolution.
LitCurate: A Configuration-Driven AI-Assisted Framework for Scientific Database Construction with an Application to Lower-Mantle Equation-of-State Data
Decades of experimental results sit in publications in forms that cannot feed data-driven modeling, and building databases from them requires both finding the relevant studies and extracting quantities with enough scientific context to stay usable. LitCurate is an open-source framework that runs literature discovery, relevance screening, full-text processing, and structured extraction with large language models as separate auditable stages, retaining intermediate results and provenance so researchers can inspect and revise any step rather than trusting a black box. Applied to high-pressure mineral physics, it produced an equation-of-state database of 1,334 entries drawn from 205 papers, linking parameters to mineral phases, compositions, equation formulations, and methods, and labeling each value as source-reported or citation-reported.
RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing
Retraining a classifier can break inputs the previous version got right, and confirming those regressions is expensive when ground truth needs human annotation or simulation, so only a fraction of inputs can be checked. RiskBlend ranks candidates by blending four signals — historical failure patterns, prediction shift, decision-boundary shift and neighborhood change — with weights learned on validation data under an APFD-squared objective, instead of relying on a single model's confidence scores. Across 1,200 configurations spanning four datasets, five classifiers and four update scenarios, it achieved the highest average APFD in all 80 dataset-classifier-scenario combinations, improving by up to 0.32 APFD over the strongest baseline; confidence-based ranking remained competitive only for linear classifiers on sparse categorical features.
Diffusion Distillation for Efficient Weather Ensembles
Diffusion models produce skillful weather ensembles but pay for it with costly iterative sampling. A supervised energy-distance distillation objective compresses a multi-step diffusion teacher into a single-step student by aligning student forecasts with both teacher samples and ground-truth observations. On global forecasting and typhoon-track prediction the student outperforms existing distillation methods, retains skill on extreme events, and matches or surpasses the teacher across key metrics using just one neural function evaluation per autoregressive step.
Efficient Auto-Interpretability of AI Models in Biology
Sparse autoencoders (SAEs) could turn biological foundation models into instruments for discovery, but three distinct questions — whether a latent is coherent, whether it can be described, and whether that description predicts anything — routinely get conflated. The proposed pipeline separates them: cross-seed dictionary stability decides which latents are worth investigating, an intruder-detection task asks whether a latent's activating examples share a recognizable pattern, and a final pass converts a candidate biological description into falsifiable in-silico predictions. On the Boltz-1 Pairformer trunk, stability prioritization found interpretable latents at 5.2 times lower measured cost (4.4 times fewer evaluations each) while recovering over half of them, though it appears to favor structure-related features over function-related ones.
From Uncertainty to Clinical Risk: Severity-Aware Conformal Planning for Interactive Medical Diagnosis
Interactive diagnosis systems decide turn by turn whether to ask another question or commit to an answer, but existing methods steer that choice by predictive uncertainty alone and ignore that missing a severe disease costs far more than missing a mild one. Severity-Aware Conformal Clinical Planning treats the process as a risk-sensitive sequential decision problem, maintaining separate diagnostic, safety, and masked-evidence beliefs, calibrating turn-specific conformal prediction sets and severity-weighted differential-diagnosis risk on held-out trajectories, and feeding that calibrated risk into Monte Carlo Tree Search to score long-horizon Ask and Commit paths. On DDXPlus and MediQ across several LLM backbones, it reaches more accurate diagnoses with fewer questions while reducing high-risk errors on severe cases.
Beyond Pairwise Graphs in Science: Hypergraph Adaptive Wavelet Operators for Parametric PDEs
Neural surrogates for physical simulation work best on regular grids, yet realistic geometries demand unstructured meshes, and graph-based operators represent only pairwise edges, missing the group-wise coupling among mesh cells, neighborhoods, and conservation volumes. HALO lifts the domain to a hypergraph and learns in its spectral wavelet domain, using Chebyshev polynomial wavelet filters to avoid eigendecomposition — localized spectral kernels at linear sparse-matrix cost — with trainable dyadic scales regularized toward tight-frame coverage so the frequency response adapts per equation. Across 2D and 3D benchmarks on structured and unstructured discretizations it is best or near-best against frequency-, transformer-, DeepONet-, state-space-, and graph-based baselines with stable multi-step rollouts, and stays competitive with the strongest fixed-discretization transformers on industrial aerodynamic meshes of a few hundred thousand points while remaining resolution-equivariant.
From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning
Financial question answering mixes tables, charts, and narrative text, and the standard Exact Match metric punishes answers that differ only in units or formatting, distorting how model reasoning is assessed. The pipeline generates synthetic question-answer pairs behind aggressive validation checks, fine-tunes smaller language models with Quantized Low-Rank Adaptation (QLoRA), and scores answers by evaluating the arithmetic expression a model produces rather than string-matching the ground truth, with a modified loss that blends cross-entropy, semantic similarity between predicted and reference expressions, and the new metric. On ConvFinQA the combination of synthetic data and the modified loss produces significant gains in question-answering accuracy over the untuned baseline.
Antipatterns in AI-assisted Qualitative Data Analysis: A Catalog of Temptations and Pitfalls for Software Engineering Researchers
Software engineering researchers are increasingly using language models for qualitative data analysis (QDA) — coding interviews, open-ended survey responses, and similar material — with little strategic guidance on what can go methodologically wrong. Drawing on the authors' qualitative research experience, the catalog names practices that look advantageous but undermine analytical rigor and validity, grouped by escalating impact into Dangerous Drivers, Operational Missteps, and Analytical Failures. The intent is both to give researchers a list of temptations to avoid and to give reviewers shared vocabulary for identifying failed practice.
DisCTI: Who Needs to Know Timely? Automated Sector-Aware Cyber Threat Intelligence Dissemination
Cyber threat intelligence (CTI) is only useful if it reaches the right industry sector quickly, yet on the Malware Information Sharing Platform (MISP) 98% of events carry no sector tag at all, forcing analysts to manually sift heterogeneous feeds. The authors recast sector-targeted dissemination as multilabel classification, build a dataset of 872 sector-labelled CTI events drawn from a threat intelligence platform, and fine-tune BERT over events expressed in the Structured Threat Information Expression (STIX) format for cross-platform portability. The resulting DisCTI classifier reaches a macro-averaged F1 of 0.89 with a Hamming loss of 0.055, meaning 94.5% of individual sector-label assignments are correct.
Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code
Prior studies count vulnerabilities in model-written Infrastructure-as-Code (IaC) without a human comparison point, so they cannot say whether models are actually less secure than engineers. GenIaC-SecBench covers 100 deployment scenarios across 12 model configurations from four vendors, producing 1,196 artifacts scanned by Checkov, Trivy, and KICS, alongside 634 human-authored templates run through the same toolchain. Because vulnerability density falls sharply with artifact size (Spearman ρ = −0.55), unmatched comparisons mostly measure length; matched on declared-resource count, every model configuration lands between 3.21x and 3.87x the human vulnerability density, with the gap widest on single-resource tasks. Vendor extended-thinking APIs cut vulnerabilities by 12.0% while prompted chain-of-thought is statistically indistinguishable from plain generation, and deployability shows no correlation with security.
Explainable Uncertainty Estimation for Reliable Medical AI
Uncertainty estimates tell a clinician when a prediction may be unreliable and explainability methods say which features drove a prediction, but the two are computed separately, so nothing indicates why a given prediction is uncertain or which additional test would reduce that uncertainty. The egRUE (Expected Gradients Reconstruction Uncertainty Estimate) method folds prediction explanations into the uncertainty computation itself and decomposes the resulting uncertainty into per-feature contributions, with proved theoretical properties and experiments against existing estimators. A user study with medical experts found that egRUE's feature-level explanations improved calibrated trust over bare uncertainty scores, raising confidence in correct predictions and lowering it on incorrect ones.
Learning to Difference: Adaptive Reversible Differencing (AdaRDiff) for Time Series Forecasting
Differencing — subtracting nearby past values to strip out trend and seasonality — is a classical fix for long-horizon forecasting, but its dependence on hand-chosen orders and periods has kept it out of modern deep architectures. AdaRDiff makes the operation learnable, using trained weights over previous time steps to produce stabilized residuals for forecasting and then autoregressively restoring the removed components; the reconstruction has a closed-form convolutional expression that parallelizes on GPU for up to a 33.7× speedup over the naive recurrence, and a two-phase training schedule separates structure discovery from reconstruction learning. Across eight electricity, weather, traffic, and energy benchmarks it reaches state-of-the-art accuracy at negligible parameter cost, and as a drop-in module it improves eight different backbones, by up to 25.9% with a linear model and 18.3% with iTransformer.
Regime-Aware Portfolio Management via Retrieval-Augmented LLM-Guided Expert Switching
Because market regimes shift, no single portfolio-management strategy stays best, so this system selects among expert strategies rather than trading directly. A dual-stream variational autoencoder encodes asset-level and market-wide state into a retrieval key, a knowledge base stores past market situations together with how each expert performed in them, and an instruction-tuned large language model reasons over the retrieved evidence to pick an expert; a monotonicity result shows that adding a locally superior expert cannot degrade the switcher. Across cryptocurrency, stock, and foreign exchange markets the selector led on cumulative return and Sharpe ratio, raising stock-market cumulative return from 26% for the best fixed expert to 34% and Sharpe from 0.74 to 0.96.
Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation
E-commerce repurchase recommenders are usually built as binary classifiers answering "will this customer rebuy within W days", which forces a separate trained model for every time horizon. Evaluated on millions of customers at a large grocery platform across more than thirty ablation configurations, survival models predicting time-to-repurchase directly replace that stack: a single Accelerated Failure Time model matches or beats three per-horizon classifiers at their own horizons while using roughly three times fewer trees in total. The empirical hazard turns out to be slightly decreasing (shape parameter around 0.9), contradicting the intuition that grocery items get more likely to be rebought the longer since last purchase, and a four-parameter calibration maps survival curves to per-horizon probabilities with no cross-horizon monotonicity violations. Calibration quality varies tenfold within the same model family, so the team ships Exponential AFT (expected calibration error around 1e-4) where probabilities are consumed and Log-Normal where only ranking matters.
BEACON: Behavior-Anchored Cross-Source Knowledge Graph Construction for Cyber Threat Intelligence
Cyber threat intelligence mostly lives in unstructured reports, and automated knowledge-graph extraction so far handles one report at a time, leaving the harder cross-source problem untouched: different vendors give the same threat different names. BEACON anchors every report on attack behaviors mapped to MITRE ATT&CK, the standard catalog of attack techniques, then attaches contextual entities such as threat actors, campaigns, and indicators of compromise to those anchors so all per-report graphs share one canonical space. Extraction uses a propose-then-verify loop grounded in both report evidence and official ATT&CK definitions to suppress hallucination, and merging proceeds hierarchically from character-level and semantic similarity to overlapping technique neighborhoods, iterating as merges pool more neighborhoods. The authors release two human-annotated datasets built from 34 sources — the largest for report-level extraction (8,395 elements) and the first for cross-source consolidation (3,487) — on which BEACON beats every baseline by at least 23% and 9% respectively.
RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents
A trading policy that reacts systematically to price moves becomes predictable to other market participants, and RetailAgent tests whether language model agents show that kind of directional structure. The agent sees anonymized intraday equity price histories plus permitted state and repeatedly chooses long or flat before the next interval's return is revealed; comparing returns during long versus flat intervals on the same stock-day, after removing the overall fraction of long decisions, isolates timing skill from exposure. Timing came out persistently negative across input modality, horizon, state, and model family, and shuffling the saved action sequences largely removes the effect, indicating the alignment between actions and subsequent returns rather than a scoring artifact. Feeding the agent its own written memories increased policy persistence, and negative timing was worst on stock-days where the agent used both actions.
Learning a Size-Weight Frontier for Synthetic-Augmented Inference
Synthetic data can shore up statistical inference when real observations are scarce, but treating generated samples as if they were real introduces bias and breaks confidence-interval coverage. The proposed framework parameterizes synthetic augmentation by two knobs, how many synthetic observations to add and what weight to give them, and estimates a size-weight frontier from a population of historical related tasks: for each weight, the largest synthetic sample size at which all smaller sizes still hit the target coverage. A finite-sample coverage guarantee holds simultaneously for every configuration on or below the estimated frontier, and in experiments augmenting opinion survey data with large language model responses hit target coverage while substantially narrowing confidence intervals.
62 more specialized papers
- Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech Han Wang, Yuhu Cheng, Xuesong Wang et al.
- UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering Mohammad Arvan, Hossein Haeri, Natalie Parde et al.
- Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis Deborah Dore, Greta Damo, Elena Cabrio et al.
- Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations D. K. C. Senevirathna, A. A. E. Nanayakkara, H. M. C. K. Kulathunga et al.
- Multiscale Community-Based Fingerprinting of Signed Functional Networks Sema Athamnah, Selin Aviyente
- CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence Pratik Ghawate, Tanvi Patil
- Destroy Me: Automatic Artifact Generation for Histopathology Images Zuzanna Krawczyk-Borysiak, Adam Krawczyk, Mateusz Miller et al.
- SETU: An Agentic Ecosystem for Multilingual, Persona-Aware Communication Coaching Jonnalagadda Maruthi Tejas, Uponika Barman Roy, Tilottama Goswami et al.
- Ab initio Modeling of MoS2/Oxide Device Interfaces with Machine Learned Electronic Structures Manasa Kaniselvan, Mauro Dossena, Denghui Lu et al.
- Towards Large-Scale Heterogeneous Data Organization for Scientific Foundation Models: A Nuclear Fusion Case Study Nathaniel Chen, Kouroche Bouchiat, Peter Steiner et al.
- Physics-informed learning for the inverse problem in resonant ultrasound spectroscopy Alejandro Cubillos Mu\~noz, Manuela Rivas, Julian Rincon
- Generative AI Expands the Intellectual Reach of Course Based Undergraduate Research Experiences (CUREs) Aditi Babar, Kristin J. Davin, Alex Dornburg
- A Framework for Object-Centric Predictive Monitoring of Collaborative Processes Daniel Calegari, Andrea Delgado, Leonel Pe\~na et al.
- SafeStep: An Interactive Demonstration of Semantic Communication for Pedestrian Safety Monitoring Christian McDowell, Andrea Panebianco, Jeremiah Yang et al.
- CARDINAL Predicts Cardiovascular Risk From Non-contrast Cardiac CT Roy Gabriel, Nattakorn Kittisut, Jamshid Hassanpour et al.
- Evaluating Loss Functions in Differentiable Out-of-Domain Sound-Matching with Partial Parameter Distance Amir Salimi, Daniel Penner, Kalvin Eng et al.
- DART-FL: Burst-Aware Multitask Federated Learning under Dynamic Inference Demand at the Edge Yiming Xie, Pinrui Yu, Geng Yuan et al.
- Beyond Non-IID: Learner--Client Distribution Mismatch in Federated Learning Yiming Xie, Lili Su, Ningfang Mi
- Leveraging a Foundation Model for the EEG-Based Diagnosis of Alzheimer's Disease Maggie Lin, Chung-Lin Hou, Tzyy-Ping Jung
- Initialization Is Critical: Advancing Federated Short-Term Load Forecasting under Load Heterogeneity via Model Initialization Jianing Chen, Vajiheh Farhadi, Yan Li et al.
- How Much Can AI Understand? Toward AI-Assisted Sensemaking of Collaborative Discussion in Groups with Shared History Soobin Cho, Mark Zachry, David W. McDonald
- Auditing Generative Audio Calls for Known-Task Audio-LLM Evaluation Mengzhe Geng
- Personalized and Multi-View Representation for Federated Cold-Start Recommendation Jaehyung Lim, Wonbin Kweon, Woojoo Kim et al.
- An Empirical Evaluation of Cross-City POI Recommendation on a Large-Scale Benchmark Peibo Li, Yang Song, Hao Xue et al.
- Low-Altitude Fluid Antenna Network with Multi-Agent Reinforcement Learning Tong Zhang, Yanfei Su, Shuai Wang et al.
- PCBnet: A Dataset and Automatic Construction of SPICE Netlists from Schematic Images Zhen Huang, Yuhao Gao, Yuzhi Liu et al.
- TI$^2$PS: A Topology-Informed Inverse Design Framework for Stochastic Multicellular Pattern Formation Kenji Komiya, Andrew Kailiang Jin, Ryo Nishikimi et al.
- A Deep Learning-Based Stacking Ensemble Framework for Turbofan Engine Remaining Useful Life Prediction Limon Bin Hossain, Md. Salehin Seyam, Md Rashedul Islam et al.
- CASTANET: Causality-Aware Spatio-Temporal Adversarial Network Using Traffic Incident Effects Toshiya Kitahara, Ryu Shirakami, Koh Takeuchi et al.
- PhyMamba: Physics-Modulated Mamba for Robust Battery Health Prognostics Sara Sameer, Yunyi Zhao, Wei Zhang et al.
- Automated Analysis Framework for Multilingual Climate-Health Literature Based on Multi-Agent Large Language Model Yuze Sun, Shihui Zhang, Jiancheng Pan et al.
- PhenoIntel: A Lifecycle-Aligned Multi-Agent Web Application for Verified, Accessible Plant Phenotype Analysis Narendren S V, Soumyashree Kar
- Comparing Classical and Quantum Machine Learning for Regression in High Energy Physics Collision Data Tariq Mahmood, Zain ul Abidin, Itzel Luviano Soto et al.
- Generalized Gibbs Ensemble Weighting for Forecast Combination Prasen R. Nuthanakaluva, Nava K. Gaddam
- CheXtriev: Anatomy-Centered Representation for Case-Based Retrieval of Chest Radiographs Naren Akash, Arihanth Tadanki, Jayanthi Sivaswamy
- Conditional Diffusion Models for Energy-Efficient Driving Hemanth Neelgund Ramesh, Andr\'e Snoeck, Chyi-Fu Hong et al.
- Under-Mattress Temporal Sensing for Next-Day Agitation Risk Scoring in Dementia Wards Zhen Liu, Marta Bono, Robbe Decloedt et al.
- Gen-TAS: A Generative AI-Aided Hardware-Software Task Allocation Framework for FPGA-GPP Heterogeneous Systems Mary Kong, Yuqin Zhao, Semih Vazgecen et al.
- Empowering Local Agriculture: A Deep Learning-Powered Web System for Identifying Bangladeshi Mango Varieties Monowar Islam, Safaruzzaman Shovo
- Text Restoration of Ancient Documents with Language Models Shibingfeng Zhang, Edoardo Caraffa, Annafelicia Zuffrano et al.
- Expert Knowledge & Machine Understanding: Bridging Reactome's Ontology with LLM Semantic Embeddings Susanna Bravi, Riccardo De Luca, Rosa Sicilia et al.
- Beyond Flat Netlist: Hierarchical Graph Representation Learning for Scalable Analysis of Sequential Circuits Jingyi Zhou, Zhengyuan Shi, Jiaying Zhu et al.
- Explainable Diabetic Retinopathy Classification Using Vision Foundation Models Abhishek Verma, Anila Krishna, Abhishek Gajanan Bankar et al.
- D-TAIA: Domain-Aware LLM Adaptation for Multi-Task Predictive Process Monitoring Sjoerd van Straten, Christine Jacob, Marwan Hassani
- Efficient Online Continual Foundation Model Fine-Tuning for Predictive Process Monitoring Sjoerd van Straten, Marwan Hassani
- Spectral Features Dominate BCG Respiratory-Event Detection: A Large-Scale Patient-Independent Comparison of Feature Groups in Sleep Apnea Patients Israel Campero Jurado, Zoe Bousraou, Lara Benning et al.
- RECAST: Recent & Context-Aware Sampling for Test-Time Adaptation in Streaming Biosignals Yong-Yeon Jo, Junho Song, Joon-myoung Kwon
- Learning to Transfer Across Modes: Towards Unified Urban Mobility Forecasting Yixuan Zhao, Man Luo
- MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in Chemistry Mahdi Babaei, Xueshen Li, Yutao Kuang et al.
- BanglaMed-QA: A Question Answering System for Healthcare Support in Bangla Rowzatul Zannat, Abdullah Al Shafi, K. M. Azharul Hasan et al.
- Propagating construction-time knowledge quality into medical question answering: A framework grounded in clinical guidelines Jie Hu, Junjie Wang, Shan Lu et al.
- Real-Time Monitoring of MHD Liquid Metal Flows with Shallow Recurrent Decoders Claudio Scardino, Stefano Riva, Carolina Introini et al.
- Real-Time Musculoskeletal Surrogates for Pediatric Cerebral Palsy: a Credibility Pilot Mohammad Arif Ul Alam
- AI as Teammate: Rethinking Task Distribution in Medical Training Fendi Tsim, Alina Gutoreva, Anthony Weiss et al.
- VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings Menghan Liu, Elynn Chen
- A Unified Framework to Elicit Structured Feedback for Interpretable Multi-Trait Essay Scoring Shihang Yang, Sanwoo Lee, Ningning Zhao et al.
- SymboLLM-FE: LLM-Accelerated Symbolic Regression for Automated Feature Engineering on Tabular Data Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou et al.
- Euclidean Fourier Neural Operators Nathanael Bosch, Niklas Frederik Schmitz, Michael F. Herbst
- Real-time virtual circuits for plasma shape control via neural network emulators: experimental demonstration on MAST Upgrade Nicola C. Amorisco, Kamran Pentland, Adriano Agnello et al.
- InstructMesh: Selective Refinement of Generative 3D Models for Fabrication Faraz Faruqi, Ahmed Katary, Demircan Tas et al.
- Advancing Interaction-Sensitive Feature Selection: Novel Relief-Based Algorithms, Expanded Comparisons, and Recommendations for Biomedical Data Mining Kia Kazemi-Nia, Harsh Bandhey, Philip J. Freda et al.
- QGPINNs: A Physics-Informed Neural Network Framework for Nonlocal Differential Equations on Quantum Graphs Vaibhav Mehandiratta, Saket Ramchandra
Large Language Models 48
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model
In 2011 IBM's Watson beat human Jeopardy! champions using a curated billion-document corpus on a cluster of POWER7 servers, frozen at build time and impossible to copy; the question asked here is whether that same snapshot of testable cultural knowledge now fits in a single downloadable file. A 9 GB 4-bit quantized Qwen2.5-14B is run over all 529,939 clues from 41 broadcast seasons (1984–2025), which the authors describe as the first evaluation over the complete corpus. Under a strict forced-response protocol with exact and fuzzy matching the model answers 67.0% of all clues and exceeds 85% on factoid categories. On clues that aired after its training cutoff it holds 65% and Claude Opus 4.8 holds 95%, while Watson scores zero by construction since it could not answer outside its curated distribution.
Accelerating LLM Inference via Vector Index Based Output Embeddings
The output embedding matrix is a memory-bandwidth bottleneck during autoregressive decoding, especially for compact models carrying large multilingual vocabularies. The output projection followed by top-k token selection is recast as a maximum inner product search over token embeddings and served by an HNSW vector index, which returns a small candidate set whose logits are scattered into a sparse full-vocabulary tensor so existing decoding pipelines work unchanged. On CPU inference with Gemma 3, Llama 3.2, and Qwen 3, end-to-end batch-size-one decoding throughput improves by up to 82% for Gemma 3 270M while generation quality holds under AlpacaEval, suggesting approximate retrieval is a workable substitute for dense output projections in latency-sensitive small-batch settings.
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
LLMs now sit on every side of evaluation — as examinees scored on benchmarks, as judges of other models, and as raters of human-written content — yet standard practice does not separate how much the rater, the instrument, and the object each contribute to a score. Rasch measurement theory (RMT) is proposed as the missing tool, decomposing ordinal ratings into separable facets on a shared scale and supplying diagnostics for miscalibration and rater bias. Fitting many-facet Rasch models to annotations from nine LLMs across families and capability levels on the Measuring Hate Speech corpus shows they differ systematically from human raters in severity, item-level calibration, question-order robustness, target-identity sensitivity, and rating-scale use — all of which conventional agreement metrics would hide.
Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection
Entity disambiguation (ED) is usually modeled as one task, though it contains two — retrieving candidate entities and selecting the right one — and dual-encoder models force a single embedding space to serve both while requiring a trained retriever that must be maintained as the knowledge graph changes. Holding an LLM selection stage fixed, the study compares sparse retrieval (BM25), web knowledge-base search, and a state-of-the-art trained dense retriever across open and closed models. A fully training-free BM25 retriever paired with an LLM selector sets a new state of the art on ZELDA, raising inKB micro-F1 from 82.3 to 86.3, while a trained dense retriever reaches 88.5. Decoupling the stages also lets the system abstain when the correct entity is absent from the candidates, reaching 90.7 F1 in a setting that rewards correct abstentions.
LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation
Bayesian network structure learning (BNSL) from observational data recovers which variables are connected but often cannot determine edge direction, while LLMs carry broad causal knowledge of uneven reliability. Both sources are expressed as Probabilistic Dependency Graphs, in which every edge holds a distribution over directed, undirected, and absent states so the two can be fused by weighted averaging. Across 26 benchmark networks combining ensembles of FGES, Tabu, and PC with Gemini, Claude, and GPT over multiple prompts and seeds, a plain 50/50 fusion improves F1 over the better single source in 22 of 26 networks, a mean gain of 0.056 (p < 0.001). The roles are complementary: BNSL supplies a high-recall skeleton at 80% versus 60%, while the LLMs orient edges at 96% accuracy versus 77%.
XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering
Multilingual multi-hop question answering benchmarks typically translate an entire example into one language, which conceals failures that occur where a reasoning chain crosses a language boundary. XHotpotQA models each instance as an evidence-dependency graph and assigns languages separately to the question, bridge evidence, answer-bearing evidence, and distractors; the audited release holds 15,661 training and 7,405 validation instances with sentence-level support supervision, and 95.60% of validation items place gold paragraphs in different languages. Across three readers, full question-evidence language mismatch is associated with 10.25 to 15.79 lower Unicode-aware answer F1 and different-script evidence with deficits of 11.98 to 23.70 points, while the corresponding evidence-selector gaps are under two points — locating the weakness in reading rather than retrieval.
DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization
Language models using Gated DeltaNet or Kimi Delta Attention swap the growing key-value cache for fixed-size recurrent states, but those states are typically kept in FP32, occupying substantial GPU memory and making updates memory-bandwidth bound. Uniform post-training quantization turns out to trade poorly here — INT8 and FP8 already hurt complex reasoning, and INT4 and NVFP4 collapse accuracy — while quantization-error energy concentrates in a few channels whose relative decay strength stays stable across prompts. DAMP uses offline calibration on both error energy and decay-based persistence to keep high-risk channels at higher precision and the rest at INT8; on Qwen3.6-35B and Kimi-Linear-48B across six benchmarks it holds accuracy near the FP32 baseline at 9.9 bits per state value while cutting recurrent-state storage by 69.1% and speeding the update kernel up to 2.01x.
Trajectory-Level Speculative Decoding for Diffusion Language Models
Diffusion language models generate tokens in parallel through iterative denoising, but decoding degenerates to one token at a time when confidence is low, and standard speculative decoding does not apply because these models speculate over denoising trajectories — multi-token updates with positions and unmasking orders — rather than left-to-right token sequences. The proposed framework drafts trajectories via confidence-stratified tree exploration, verifies them with blockwise parallel evaluation under bidirectional attention masking, and adds inter-block speculation for cross-block lookahead, with a formal account of when the procedure is exact and of trajectory drift as the price of more parallelism. Built on Fast-dLLM's dual-cache infrastructure, it cuts denoising iterations by 30-40% and raises tokens per step from 2.6 to 4.3, giving 7-14x speedup over vanilla diffusion decoding and 1.3x over Fast-dLLM with under 1% accuracy change on reasoning and code benchmarks.
When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages
Subword tokenizers push the frequency statistics of dominant languages onto low-resource languages that share their script, while byte-level models avoid this but mismatch the word-level granularity that many tasks need in non-Latin scripts. Hierarchical byte models that group bytes into word-aligned chunks normally demand huge training budgets and misalign representationally against a frozen subword language model, so the authors initialize byte embeddings directly from the frozen model's subword representations, add a chunk alignment loss projecting dynamically grouped byte chunks toward precomputed subword targets, and interleave lightweight part-of-speech supervision to guide boundary detection. Across six languages the tokenizer-free approach improves word-level morphological tasks, with up to a 13.3% gain on part-of-speech tagging.
Knowing Before Answering: Decoding Language Models for Reliable RAG
Retrieval-augmented generation can hand a model documents that are insufficient or that contradict each other, and the system ideally recognizes which case it is in before answering. The authors build a controlled benchmark of fictitious retrieval contexts labeled answerable, insufficient, or conflicting, then train a lightweight linear probe on hidden activations and attention-derived features to make that three-way call. Across 16 language models of varying architecture and size, the feature-based router consistently beats prompting baselines and specialized RAG models, with the most informative signal appearing in middle layers and hidden activations outperforming attention values or MLP outputs in most models.
Below the Noise Floor: Bimodal Seed Collapse and Distinct Failure Modes in Small-Model Knowledge Distillation
Knowledge distillation gains for small models are usually reported from a single seed, at scales where nobody has measured run-to-run variance. Eight distillation variants were compared against plain supervised cross-entropy on a 740-instance healthcare API routing task with a 1.5B Qwen student and a 20B teacher, using three to six seeds for key configurations. Per-seed standard deviation ranged from 2.8 to 48.7 percentage points, wide enough to swallow every claimed gain below five points, and three of seven variants showed bimodal collapse where at least one seed in three to five lands below 55% accuracy — including a previously undocumented truncation mode in reasoning_kd that emits reasoning but stops before naming a function (0.9% accuracy). Only progressive_kd and rank_kd avoided collapse, and a cross-split +3.78 point gain from input enrichment reversed to -2.70 under controlled within-split multi-seed retesting.
Fast Weight Attention for Continual Learning
Recurrent fast-weight memories and selective state-space models squeeze an expanding context into a fixed-size state, which makes the state transition an online learning rule; under read-after-write autoregressive semantics the correct local training example at each step pairs the previous key with the current value, not the same-step key, and the common same-step association optimizes a different internal objective. From that alignment the authors derive normalized first-order updates for squared-error regression — Falcon-1 (a scalar normalized least-mean-squares rule), Falcon-2 (per-column) and Falcon-3 (sliding-window mini-batch) — plus inner-product variants, each with recurrent, masked-parallel and chunk-parallel forms and a numerically stable positive-decay renormalization. Representative variants stay competitive on language modeling and improve length extrapolation on variable-digit addition, with the framework separating temporal alignment, plasticity, forgetting and bounded rehearsal into independent knobs.
KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution Preservation
Fine-tuning-based knowledge editing normally optimizes cross-entropy, which raises the probability of the edited answer without constraining how the rest of the output distribution shifts; across sequential edits that unconstrained redistribution accumulates and degrades locality, meaning the model's behavior on unrelated inputs changes too. KLOD replaces this with a bounded objective that stops amplifying the target once its probability crosses a threshold, while preserving the target-excluded distribution at target positions and the full next-token distribution at prefix positions. On CounterFact and ZsRE with Llama3-8B-Instruct and Qwen2.5-7B-Instruct, it substantially mitigates locality degradation while maintaining high edit reliability, and the probability threshold exposes a tunable generalization–locality trade-off.
AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not
Machine-generated text is known to be stylometrically distinguishable from human writing, but models are now used just as often to edit human drafts, and it was unclear whether editing leaves the same trace. Measuring stylometric features across 8 large language models and 5 domains, the authors find generation leaves a consistent footprint driven mainly by entropy and lexical diversity, while the remaining features vary by domain and generator. Editing does not reproduce that footprint: edited text shows a small rise in lexical diversity alongside a drop in entropy rather than the joint increase typical of generation, with lexical density becoming the dominant signal, so stylometry separates edited from generated text but struggles to separate edited from human text.
SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models
Spiking neural networks (SNNs) promise energy-efficient language modeling via sparse, event-driven computation, but are hard to train from scratch, so practitioners distill from a pretrained artificial neural network teacher — using fixed corpus prefixes, even though inference conditions on the model's own generations. That prefix-source mismatch shows up as both output-policy divergence from the teacher and drift in internal spiking dynamics, and a naive on-policy distillation variant turns out to suffer delayed rollout-feedback collapse. SpikeOPD stabilizes on-policy training by combining full-KL teacher correction with policy anchoring to a frozen reference SNN on matched prefixes and layerwise spike regularization, improving average accuracy over standard distilled SNNs by 0.8, 1.7, and 2.9 points at 0.125B, 0.35B, and 1.3B parameters while preserving sparse compute.
HyQuant: Hybrid-Precision Quantization for LLM Attention
Quantizing the attention module of a large language model to very low bit-widths introduces errors large enough to hurt accuracy, and prior work mostly attacks this with outlier-smoothing transformations. HyQuant instead keeps a small accuracy-critical subset in high precision — the vertical-line tokens that attention repeatedly attends to, plus a local sliding window — while quantizing everything else, using lightweight attention-pattern signals to pick that subset. The same principle applies in both phases: a hybrid-precision attention operator during prefill, and KV-cache compression with dequantization fused into the attention kernel during decode, yielding near-lossless accuracy across models and tasks from a simple design.
Entity-Memory Graph Retrieval Improves Evidence Coverage in Long-Conversation Question Answering
Long multi-session dialogue strains retrieval, because ranking by dense cosine similarity can drop a neighbouring turn that supplies necessary context. The retriever described keeps dialogue turns as verbatim memory nodes, links repeated mentions through shared entity nodes, connects adjacent memories with directed chronological edges, and answers queries by entity gating, semantic fusion, one-hop chronological recovery, then dense backfill — compared against a dense control matched on vectors, context budget, answer protocol, and evaluator to isolate the effect of graph structure. On 1,986 questions from ten LoCoMo conversations, evidence recall at top-25 rises from 79.7% to 84.5%, with the advantage holding from top-5 through top-50; notably, no matched cutoff supports any difference in final-answer F1, so the demonstrated gain is in retrieval coverage alone.
When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation
On-policy distillation has the teacher score the student's own generated prefixes, but that guidance is not always trustworthy: it can discourage moves toward correct trajectories or push the student toward incorrect ones, working against the outcome reward the training is supposed to maximize. Reward-Aligned On-Policy Distillation (RA-OPD) checks, for each sampled trajectory, whether its trajectory-level distillation return agrees in sign with its outcome reward, and drops the trajectories where the two conflict — a filter that costs no extra compute. Across seven math and three code benchmarks with Qwen3 and DeepSeek-R1 family models, it significantly outperforms standard on-policy distillation and the other variants tested.
QUORUM: QUality-Optimized Routing Using Multiple annotators
Large language models can annotate data far more cheaply than humans but their reliability varies sharply per instance, handling easy inputs well while failing on examples needing nuanced reasoning or context. QUORUM is a budget-aware routing framework that decides per instance whether to send it to a human or an LLM annotator under a fixed budget, estimating difficulty from feature-based signals instead of model confidence or uncertainty, and allowing several annotations per instance that are combined through agreement-based rewards. Across closed- and open-ended annotation tasks in English and multilingual settings, it improves annotation quality by up to 34.4% while cutting cost by 8.8% relative to competing routing methods, with code released publicly.
A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint
Rather than quantizing every layer uniformly as GPTQ or AWQ do, or allocating bit-widths without demonstrating actual speedup as MixLLM and TorchAO do, this work frames per-layer bit allocation for Gemma-3-1B as maximizing latency reduction subject to a budget on generation-quality loss, reusing a layer sensitivity profile from the authors' earlier SA-PTQ work through TensorRT-LLM's activation pass-through mode. Measuring 13 W8A8 variants on an RTX 5090 across block groupings (5+5, 10+10, all 26 layers), they find integer arithmetic pays for the quantize/dequantize overhead in feed-forward networks and the language-model head, but not in attention at short context lengths, where the extra step is a net slowdown; a manual SmoothQuant implementation was needed because export failed. The best configuration under minimal degradation, feed-forward 5+5 plus the head, gives an 11.0% latency reduction at 98.90% Top-1 agreement and +0.85% perplexity, rising to 19.1% speedup with looser quality tolerance.
Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning
Knowledge-intensive question answering requires models to ground answers in supplied evidence, but entity mentions in context can trigger memorised associations that yield plausible answers unsupported by that evidence, and existing abstention methods based on uncertainty scores or evidence-sufficiency checks never test grounding directly. Twin Worlds instead builds several parallel versions of an input through typed entity substitutions that preserve the relational structure while stripping away the model's parametric priors, then checks equivariance: a grounded answer should shift in step with the substituted entities rather than staying fixed. Violations of that expected correspondence are used as the abstention signal, and across four benchmarks and three model backbones the method identifies ungrounded answers more reliably than uncertainty- and sufficiency-based baselines.
Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms
Inference services bill per token while GPUs consume energy over whole inference windows, an accounting mismatch that lets average per-token energy fall even as total request energy climbs. A decomposed model splits consumption into a one-time prefill plus generation setup cost and a marginal per-step cost, measured on NVIDIA H100 and H200 GPUs across dense and mixture-of-experts models as a function of batch size, context length, and output length. For Llama-3.2-1B on an H200 at batch-16 and 4K context, stretching output from 10 to 512 tokens drops token energy from 7.46 to 0.72 J/token while total window energy rises from 1.19 to 5.93 kJ; batching also lowers token energy but the benefit shrinks with longer context (6.31x at 512-token context versus 1.17x at 4K), and sparse expert routing inflates fixed energy at low concurrency in a way batching largely erases.
SEPO: Evidence-Grounded Prompt Optimization via Structural Editing
Automatic prompt optimizers that only call an API are usually called interpretable, but each iteration still rewrites the whole prompt as one opaque string, so the record left behind is a series of full-prompt diffs rather than edits anyone can localise or reason about. SEPO (Structural, Evidence-grounded Prompt Optimization) runs multiple search trajectories that make local edits to stable, typed units within a two-layer prompt schema, links each edit's intended and realised structural operation to the specific examples it newly fixes or breaks, and carries that edit-effect record forward to inform later proposals on the same branch. On a 14-task held-out suite it beats GEPA by 3.1 points on Llama-3.1-8B-Instruct and 2.2 points on Qwen3-8B, reaching 61.9% and 73.3% macro accuracy while spending 2.9M optimisation tokens against GEPA's 4.1M and producing prompts more than five times shorter.
H-Scale: Hessian-Guided Scale Refinement for NVFP4 Sub-Byte LLM Inference
NVIDIA's Blackwell architecture natively supports NVFP4, whose micro-blocks of 16 weights capture local distributions and isolate outliers well but create a large, sensitive space of per-group scaling factors that existing post-training quantization work has mostly ignored in favour of refining the quantized weights themselves. H-Scale is a lightweight post-processing step that picks hardware-valid group scales using a diagonal second-order proxy derived from calibration activations, so the objective targets layer output perturbation rather than plain weight reconstruction error. It drops into existing NVFP4 pipelines as a replacement for round-to-nearest scale selection, needs only modest offline calibration, and adds zero inference-time overhead while improving a broad range of NVFP4 baselines and moving several variants closer to the BF16 reference.
The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues
Computational work on social power in dialogue has been confined to narrow linguistic and cultural settings, with datasets lacking the demographic and relational detail needed for cross-cultural comparison. The authors pair a schema grounded in social science theory with a native-speaker annotation pipeline and a cross-lingual analysis interface, producing a corpus of 15,836 annotated instances from 100 scenes in French and Egyptian Arabic films. Annotators agree strongly on observable demographic and contextual attributes but diverge on interpretive ones such as power asymmetry and intention alignment, and an evaluation of six large language models and multimodal LLMs finds persistent gaps between human and model agreement on relational and theory-of-mind reasoning.
Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result
Because a byte-level byte-pair-encoding tokenizer is an ordered list of merge rules, applying only a prefix yields a nested vocabulary whose token ids are the first rows of the full one — so a single model could in principle serve several vocabulary sizes and be deployed at any of them by slicing its embedding and output head. The authors pre-registered five claims with margins, seeds, and a stop rule, then trained 30 models with 3.1M- and 10.6M-parameter bodies on 200M tokens each. Slicing is bit-for-bit exact across 76 checks and removes 66% of deployed weights with no latency change, but the shared model trails a fixed-vocabulary specialist by 3.64% bits per byte at 32k against a 1% pre-registered margin, with a 2×2 ablation attributing the cost to output restriction rather than the control token; multi-cap training does buy robustness, degrading 12.5–15.4 points less under typographical noise.
FinExam-10K: When Retrieval Helps Financial Reasoning?
Professional finance exams demand domain knowledge, calculation, and judgment together, yet no benchmark had covered the full structure of the Chartered Financial Analyst (CFA) and Financial Risk Manager (FRM) programs under one protocol. FinExam-10K contributes 10,198 expert-reannotated questions across CFA Levels I–III and FRM Parts I–II, half released and half sequestered for a maintained leaderboard, split into a full-coverage track and a context-complete track where the supplied record suffices to answer. Across 17 models the best overall accuracy is 85.29% but drops to 34.68% on the frozen hard band, and all 17 fail the same 47 context-complete items; retrieval augmentation rescues hundreds of errors while overturning as many correct answers, for little or negative net gain, and a learned gate that fires retrieval on 7.9% of questions lifts held-out accuracy only from 70.83% to 71.23%.
Stay Within Your Bounds: Distance-Guided Decoding for Guaranteed Context-Free Grammar Compliance
Grammar-constrained decoding keeps language model output syntactically valid for formats like JSON, SQL, and code, but typical decoders only enforce local prefix feasibility, so a prefix that is extendable in principle can still fail to reach acceptance under tokenizer-grammar mismatch and a finite token budget. The proposed decoder precomputes bounded pushdown-automaton summaries offline, labeling reachability and an upper bound on distance to acceptance, then uses those estimates online for horizon-aware pruning and beam search. Every output is guaranteed to be accepted by the target context-free grammar, and experiments on JSON, SQL, and Linear Temporal Logic (LTL) report both consistent validity and better completion quality than existing baselines.
Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation
Agents that emit JSON, SQL, or function calls fail outright when one field is wrong, and constrained decoding already tracks parser transitions that reveal which tokens carry schema-critical decisions — a signal current key-value (KV) cache compression ignores. PASK (Parser-Aware Structural KV Persistence) converts those parser states into layer-group-specific decisions about which KV entries to keep, using task-error sensitivity to set minimum protection floors and attention-output distortion to allocate what capacity remains; an offline calibration step compiles this into a policy so only a lightweight lookup runs at inference. At a total KV budget of 0.33 on Qwen3-4B, it beat the strongest compressed baseline by 17.39 percentage points on average across eight BFCL non-live and Live subcategories, and in serving reached up to 2.2x higher throughput and 3.3x lower time-per-output-token at roughly half the peak GPU memory of a full cache.
A Probabilistic Interpretation of KV Cache Eviction
Dropping entries from the key-value (KV) cache buys throughput at supposedly negligible quality cost, but the selection rules in the literature are largely heuristic and the problem itself has never been stated formally. A probabilistic formalization shows the exact eviction problem is computationally hard, then reframes it as expectation estimation that can be approximated by sampling — which also makes it feasible to correct for evicted entries during decoding, something prior work ignored. Under this view existing eviction methods are zero-variance biased estimators, and adapting them to add decode-time correction gave more robust behavior across tasks at competitive performance for the same compression budget.
Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
Pretraining dense large language models on English-prevalent corpora, this study maps how optimal learning rate and batch size scale jointly and how each evolves marginally with model capacity and data budget, fitting a model of those relationships. Using a Warmup-Stable-Decay schedule, it also measures what learning-rate annealing buys across a wide sweep of settings and whether hyperparameters chosen for the stable phase transfer to the decay phase, then evaluates recently proposed loss scaling forms that explicitly model the capacity-data interaction. Those forms captured both undertraining and overtraining regimes well across the experiments, and the complete collection of pretraining runs is released open-source as a baseline for future OpenEuroLLM models.
When Linguistic and Internal Confidence Diverge in Large Language Models
Asking a language model how confident it is only helps if that spoken confidence tracks the model's internal uncertainty, which this study tests across 8 classification tasks, 2 generation tasks, and 30 models from three families, comparing verbalized confidence against logit-based confidence and semantic entropy. Instance-level association between stated and internal confidence is weak on average, improving only on easier items and stronger base models; instruction-tuned models report higher confidence and sometimes correlate better but show larger confidence gaps and worse calibration. Prompt changes mostly shift the distribution of reported numbers rather than the underlying alignment, with attitude cues inflating confidence for nothing and score exemplars preserving rank-order signal only when they avoid collapsing onto a few values. The authors describe verbal confidence as a lossy channel that can carry useful ranking information without being calibrated, and argue it needs multi-axis diagnostics before use in reliability pipelines.
CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
Cultural evaluation of language models is usually multiple-choice factual recall, which misses the common case of a user asking for practical help over several turns in a culturally specific situation. CultureConverse is a simulation and evaluation harness covering 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains, where an assistant must infer cultural constraints from partial information over a scored multi-turn episode; the released CultureConverse-DS dataset holds 14,610 benchmark episodes and 274,295 oracle-guided dialogues. Across 18 evaluated models GPT-5 mini scored highest on assistance quality, human annotation supports the automatic judge as a proxy, and fine-tuning on 27,860 high-quality samples improved in-domain assistance while transferring to out-of-domain cultural multiple-choice and safety classification benchmarks.
Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL
In-context learning (ICL) text-to-SQL systems keep stacking modules around the base generator, yet papers report only aggregate end-to-end accuracy, leaving the marginal accuracy and cost of each design choice unquantified. The authors implement 17 paradigm-level configurations spanning five recurring pipeline modules in one controlled codebase and attribute each one's contribution and token cost across four backbones of differing capability and reasoning style. Execution-feedback refinement is the only paradigm whose benefit holds universally and at consistently low cost, most other modules help only under backbone-dependent conditions, input token demand tracks pipeline structure while output demand tracks backbone generation behavior, and a fixed budget is often better spent on a more elaborate pipeline over a mid-tier backbone than on a frontier model with a lean one; the resulting tiered guideline transfers to five additional backbones without re-running the search.
Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining
Pretraining dominates LLM training compute, yet noise-dominated gradients and an ill-conditioned loss landscape leave progress slow along flat directions — the small-eigenvalue directions that drive most of the final loss reduction — which adaptive optimizers like AdamW and Muon mitigate only indirectly through gradient normalization. The proposed optimizer applies multiscale momentum solely along flat directions, pairing a slow-decay component for noise reduction with a fast-decay component for rapid curvature adaptation, and adds a sphere constraint to prevent the parameter inflation and overly fast effective-learning-rate decay a naive combination would cause. Experiments show it significantly accelerates Muon across dense and mixture-of-experts architectures at 0.12B to 2.3B parameters, with theoretical analysis supporting the flat-direction design.
Sliding-window beats linear attention
Retrofitting pretrained LLMs with linear attention has been promoted as the fix for quadratic attention's ever-growing key-value cache, but the authors argue this line of work has never been compared against simple baselines. Testing sliding window attention (SWA) with attention sinks against post-trained linear-attention models across several LLMs and downstream tasks, they find SWA matches or beats them, and on long-context tasks (Needle-in-a-Haystack and BABILong) scores 2 to 10 times higher. Since SWA requires no post-training and is fast and memory-light, they recommend it over converting models to linear attention, which they suspect would need training from scratch or extensive post-training merely to match.
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
Factual question answering benchmarks assume a single canonical answer, which hides whether a model retains genuinely divergent accounts of long-tail facts. ElephantBench is a closed-book probe of 1,094 questions built by an auditable graph-based pipeline that retrieves related documents from a low-exposure web corpus, identifies naturally occurring disagreements, converts them into multi-account records, and verifies each answer against the source documents, authoritative public sources, and human annotators. Across 32 models, the strongest recovers both accounts on only 52.4% of questions, recalling one and omitting the other on nearly all the rest; larger models and inference-time reasoning improve recall without eliminating the incompleteness, and corpus analysis links completeness to how much exposure the minority account has.
How Proper Scoring Rules Shape LLM Forecasting
Five proper scoring rules are compared as reinforcement training objectives for LLM forecasters predicting binary outcomes of already-resolved real-world events. Although all five theoretically incentivize truthful probability reporting, the resulting models differ substantially in calibration, probability usage, and their decomposition into bias, information, and noise, even when aggregate accuracy is similar — the Brier-trained model wins on Brier score and AUC-ROC while the log-score-trained model wins on log score and calibration error. The takeaway is that reward choice shapes the structure of forecasting errors, not just their magnitude, though each condition used a single training seed so some gaps may be stochastic.
Ladders in Chaos: When, How, (and Perhaps Why) Does Test-Time Scaling Improve LLM Machine Translation
Test-time scaling comes in two flavors — sequential, where each attempt conditions on the previous one, and parallel, such as independent sampling with reranking — and their relative merits for machine translation had not been characterized. Across sampling budgets, sequential sampling reaches a higher performance ceiling and yields a more diverse candidate pool, especially at small budgets, and multidimensional human analysis of Best-of-N outputs shows it mainly improves fluency and naturalness while sometimes degrading accuracy at large budgets. Controlled experiments attribute part of the gain to the model seeing more target-side context, with ablations showing robustness to temperature but sensitivity to how that context is assembled.
Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration
Training Mixture-of-Experts (MoE) language models with expert parallelism spends a large share of wall-clock time on all-to-all token dispatch and combine collectives. The communication-efficient variant CE-MoE decouples token-mixing depth from channel-mixing depth: instead of placing an MoE layer after every attention or Mamba-2 block, it concentrates expert capacity in a few routed MoE layers and restores depth with extra token-mixing and dense feed-forward layers. Over a scaling ladder from 2B to 31.5B total parameters with total and activated parameters matched, CE-MoE matches full-MoE validation loss and downstream scores while cutting training cost, and at 31.5B it uses 33.3% fewer GPU-hours while improving average downstream score and inference throughput.
DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging
Merging several task-specific fine-tuned LLMs into one multi-task model without retraining introduces representation bias, a systematic drift between the merged model's hidden states and those of each source model; prior correction work targeted encoder-based vision models. Two complications specific to decoders are identified: the causal attention mask lets bias accumulate across token positions, requiring position-dependent correction, and high-entropy decision-critical positions matter far more than low-entropy ones. DARTS addresses both with an entropy-weighted L1 loss that upweights correction where errors most affect generation, plus a per-position additive bias term. On Llama-2-7B merges evaluated across HumanEval, GSM8K, and AlpacaEval, it improves substantially over standard surgery while adding only 0.1% extra parameters.
7 more specialized papers
- Informational Antilocality and the Locality Bias in LLMs Andrew McInnerney, Shane Storks, Steven Abney et al.
- Representation of syntax in LLMs through the lens of linear distance and similarity-aware entropy Juan Pablo Vigneaux, Mary Kennedy, Khalil Iskarous et al.
- PersonaEdit: Representative Sample Selection for Personalized Model Editing You-Mei Huang, Chung-Chi Chen, An-Zi Yen
- SimpCue: Cue-Based Prompting for Multilingual Text Simplification Mehrzad Tareh, Horacio Saggion, Stefan Bott
- CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms Kaiyan Zhao, Zhongtao Miao, Zheyong Xie et al.
- Embedding Models for Stance-Aware Argument Retrieval Angelo Sparacino, Francesca Toni, Adam Dejl
- Stranger, Fan, or Peer? A Systematic Study on the Role of Interlocutor in Persona-Based Dialogue Generation Daniela Occhipinti, Malvina Nissim, Marco Guerini
Agents 46
PACE: Publisher-Adaptive Content Extraction via Agentic Automation
Web content extraction for LLM data pipelines forces a trade among general extractors that break on publisher-specific layouts, direct LLM extraction that is costly and slow at scale, and hand-written parsers that require ongoing human maintenance. PACE uses LLMs offline to analyze representative pages from a publisher and aggregate reusable extraction patterns into a configuration; at inference time that configuration instantiates a fixed deterministic extractor template, so extraction itself makes no LLM calls. Across article-body, metadata, and multimodal targets including images and tables, it beats scalable non-manual baselines and approaches the quality of manually engineered publisher-specific parsers.
Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields
Recovering a partial differential equation (PDE) in heterogeneous media means identifying the governing operator and the unknown spatial fields that parameterize it at once, and the two are entangled — a sufficiently flexible field can conceal a structurally wrong law when only one trajectory is available. HER-PDE runs a hypothesize-evaluate-refine agent loop that reads two noisy trajectories from different excitations, proposes complete expression-tree hypotheses, and scores them through an evaluation interface that estimates only the fields a hypothesis explicitly declares, never silently adding terms, ranking candidates by bidirectional cross-excitation transfer before auditing the winner on a sealed time interval. Across five two-dimensional systems observed with 5% relative Gaussian state noise it recovered the generating operator in all five cases, including equivalent signed-field and product-rule parameterizations, with the nine unknown coefficient fields reaching median Pearson correlation near 0.85 and median relative L2 error near 0.28.
Benchmarking General Mobile Assistants in Challenging Real-World Scenarios
Existing mobile agent benchmarks such as AndroidWorld and MobileWorld cover a narrower slice of app use than real users exercise, so GMA adds seven applications built from open-source projects spanning domains like lifestyle sharing and travel planning, plus 300 tasks across four difficulty tiers from atomic actions to long multi-step workflows. Across eight frontier models, performance declines substantially as task complexity increases, leaving current agents well short of handling realistic user requests reliably. Controlled ablations of agent harness choices — context retention and explicit state tracking — run under a shared environment, model setting, and task taxonomy show harness design meaningfully helps on demanding workflows, though which specific design helps varies by foundation model.
Thinking Costs Tokens: When More Structure is Worth the Price
Adding search, verification, and revision scaffolding to a language model consumes the same token budget it is meant to spend wisely, raising the question of where the break-even point lies. Two systems were compared on the FinQA and TAT-QA financial reasoning benchmarks using GPT-5.4 mini across 14 budget tiers from 250 to 42,000 output-equivalent tokens and 1,000 cases each: a single-call monolith versus a verified search architecture with planning, label-blind checking, and repair. At 1,000 tokens the monolith reaches 18% accuracy while verified search scores near zero because planning overhead leaves no room for an answer, but from 1,500 tokens onward verified search leads, topping out around 44% versus 40%. The crossover falls between 1,000 and 1,500 output-equivalent tokens, confirmed by an intersection-union test at p ≤ 0.001.
WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning
Reinforcement learning for mobile graphical user interface (GUI) agents normally requires large volumes of live Android interaction, which is slow, expensive, and unstable. WM-R1 replaces the real environment with a world model that supplies all state transitions during rollouts, and additionally embeds the world model in the agent's chain of thought so it can simulate the consequences of candidate actions before committing. The design permits massively parallel, step-level trajectory generation and uses a multi-dimensional rule-based reward covering task success, trajectory efficiency, and world model usage, trained on a curated set of 2,000 hard tasks. On Android benchmarks the resulting agents outperform both GRPO-only baselines and inference-time simulation methods.
If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary
An agent handed a human's credential inherits that person's full reach without the judgment that normally constrains it, and every over-broad query it makes still looks credential-valid to the backend. Out-of-Band Policy Enforcement (OBPE) puts a trusted boundary outside agent reasoning that authorizes the typed operation and resource, narrows the query before the backend call, then filters records and fields or masks values on the way back, with semantic gating able to deny or hold a call based on argument values or external state; the authors prove the policy plan is order-independent and that agent policy can only narrow, never widen, the data owner's ceiling. Benchmarking prompted agents against Jira and ServiceNow mocks across four models with 20 adaptive red-team tasks, trace failures fell from 57.6% to 0.2% over 3,621 trials while task fulfillment dropped from 79.1% to 60.9% and paired safe-and-useful completion rose 21.8 points. They release an HTTP proxy prototype with a typed Cedar policy core, and note residual leaks where answers reconstruct values that never entered context or use filtered row counts as an oracle.
First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents
Qwen-GuidePlay-2B is a 2-billion-parameter model for dialogue-game interaction, fine-tuned from Qwen3.5-2B in three stages: supervised fine-tuning on only successful Playpen trajectories, then weighted turn-level fine-tuning, then teacher-guided fine-tuning where a larger model fixes formatting and scores examples but never writes new gold actions. It reaches 57.12 clemscore on the public validation set and the second-highest clemscore delta among challenge submissions, about +36 over its base model. Imitating whole trajectories appears to buy basic playability while turn-level and teacher-guided training improve decision quality, and more elaborate procedures such as replay-repair and hard-example mining did not help, suggesting careful data curation matters more than aggressive training changes at this scale.
Agents for Everyone: A Workshop Framework for Building Agentic AI Capabilities in a Distributed Curation Community
Biological database curators could benefit from agentic AI but are blocked by practical barriers — no access to agents, no local tooling, no training. The Gene Ontology Consortium addressed this with a cloud environment built on JupyterHub using Claude Code as a universal harness, letting curators drive an agent session from a browser terminal through a single shared API gateway with no subscriptions or local installs, plus four training modules progressing from basic tool use to pathway curation with the existing GO-CAM (Gene Ontology Causal Activity Model) tool. Thirty-seven participants completed the four-hour workshop, and the authors conclude that building agentic capability in a distributed scientific community is mainly a matter of removing access barriers, designing workflows, introducing capabilities gradually, and giving curators hands-on practice evaluating agent output.
PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation
Estimating a product carbon footprint (PCF) — the greenhouse-gas emissions attributable to a physical product — is increasingly handed to LLM agents, but scoring only the final total lets errors cancel out and hides where the reasoning broke, while scoring subtasks in isolation misses compositional effects. PCFBench carves the workflow into six independently gradable tasks spanning decomposition, retrieval, ontology matching and numerical extraction, with 614 expert-labelled items probing under-specification, conflicting context and numerical constraints. Across eight frontier LLMs from four providers no model dominated: the strongest came within a factor of two of declared totals on 77% of products, but that rate dropped to 37–58% when the footprint was built up step by step, with only 45–75% of answers respecting mass conservation.
The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs
Hidden states carry information about model behavior that is hard to extract from inputs and outputs alone, which matters once a model's tool calls reach external systems. Linear probes trained on those hidden states were evaluated for detecting incorrect tool calls across 18 tool-calling LLMs on the Berkeley Function Calling Leaderboard. Probing caught a range of error types, including arguments with the correct type but the wrong value, which standard logging frameworks would not record, with effectiveness depending on model size, probe layer and post-training type; probes also generalized to error types absent from their training data.
Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models
A tool-equipped language model can commit to a final claim its evidence does not support even when one available call would settle the question and its instructions explicitly forbid guessing. The failure is split into two measurable quantities: occurrence, judged from visible evidence and the final claim without consulting the answer key, and conditional repair, measured by replaying each case from an exact copy of its state with alternative tool responses matched in structure and length that differ by a single character. On a fixed Qwen3-32B setup, 33 of 512 first responses ended in an unsupported claim, and resolving evidence repaired all 33 while a matched but uninformative response repaired none; an automatic checking rule added 21 evidence calls over 64 cases, corrected all 10 wrong claims and never turned a correct answer wrong, whereas a Gemma 4 setup always called the tool and produced no unsupported claims.
Credo: Reusable Declarative Primitives for Agentic Workflows
An LLM application depends as much on its harness, the program deciding what each call sees, how many calls to make, and which answers to trust, as on the model, and coding agents that search for good harnesses emit opaque imperative code whose logical steps, runtime signals, execution decisions, and prompt strategies stay implicit, forcing every new task to restart the search. Credo recovers a structured declarative description from a searched harness, tags each extracted primitive with metadata, and catalogues everything with provenance so that a compiler can bind stored primitives into harnesses for new tasks without searching from scratch. The paper reports preliminary results and lays out a research agenda aimed at the database community, including cost-based compilation over declarative catalogs and catalog maintenance under model and workload drift.
ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQL
Reinforcement learning from execution feedback has pushed text-to-SQL accuracy up, but most approaches treat SQL generation as a single turn, leaving no room to recover from errors. ReToolSQL pairs a supervised warm start on rejection-sampled reasoning traces with agentic reinforcement fine-tuning over multi-turn tool-use trajectories, on the argument that the two act on complementary axes: supervised fine-tuning expands which questions are solvable at all, raising pass@k coverage on the hardest cases, while reinforcement fine-tuning converts that coverage into single-pass accuracy by teaching the model when to verify, what evidence to retrieve, and how to repair faulty SQL. Applied to a 31B instruction-tuned Gemma 4, the combined pipeline reaches 74.32% execution accuracy on the BIRD-SQL development benchmark single-pass and 74.77% with self-consistency, reported as first on the single-model development leaderboard at the time of writing, using composite rewards anchored on execution correctness and no human annotation beyond the benchmark.
CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action
Instructions given to embodied agents in natural language usually carry constraints that must hold as the world changes, and the free-form programs an LLM writes offer no stable object to verify, compose with new constraints, or repair from a failing trace. CEDAR grounds instructions as regular languages over environment event traces, using a language model for semantic judgments and execution traces for counterexample-guided correction, and represents both learned skills and specifications as deterministic finite automata, so a skill can be intersected with a constraint such as sleeping at night or staying in one biome to yield a controller that enforces the constraint by construction rather than by repeated prompting. In Minecraft, given the same simulator observations as a program-generating baseline, it maintains temporal and spatial constraints the baseline fails to preserve while amortizing skill reuse and reducing cumulative LLM queries.
CURA: Certified Runtime Alarms for Computer-Use Agents
Self-report is the cheapest oversight channel a deployer has over computer-use agents, and it fails exactly where oversight matters: across 361 OSWorld tasks, a pipeline of read-only feasibility gate, planner, and GUI executor reaches a mean score of 82.9 against a 72.4 human reference, yet 64 of its 71 failures end with a success claim and the explicit failure affordance is never used in roughly 9,100 calls. CURA is an external monitor that reads only harness-visible telemetry, with no model internals, extra model calls, or prompt changes, turning the running trajectory into a sequential test with certified false-alarm control. At a 0.10 alarm level its CUSUM detector catches 42.3% of failures a median of 31 steps before termination at a realized false-alarm rate of 0.066; retrospectively its margin over a total-token baseline is not statistically significant, but online it recalls more at matched certified budgets, and alarm-gated escalation to a frontier overseer recovers 23 of 70 failures for a mean score of 86.8.
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
Coding agents are graded on the SWE-bench family, whose tasks come from curated GitHub issues that are long, structured, and information-rich, whereas real user requests are short and unstructured. Applying a six-category information taxonomy and four dimensions of linguistic style to real prompts from SWE-chat versus problems from SWE-bench Verified and Pro, the authors find that requests carrying only a problem statement account for 88% of real prompts but 7% of benchmark problems, and 87% of real prompts are casually written against 94% formal benchmark problems. RealSWE rebuilds 381 task families whose variants share a task and gold patch but differ in information composition and style, and across seven contemporary models realistic inputs lower resolution rates by 6.4 percentage points on average and can change model rankings. Controlled analysis shows that including desired behavior and motivation significantly affects performance while environment information and reproduction steps merely add tokens, and that style has only small model-dependent effects, so users get a concrete instruction: say what you want to happen and why.
FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling
Autonomous agents that build predictive models from electronic health records (EHRs) are limited to one hospital's data and tooling, and patient privacy blocks direct collaboration, while conventional federated learning only shares model parameters and discards the modeling know-how an agent accumulates. FedEHR-Agents federates that experience instead: each hospital's agent handles preprocessing and model development and refines local modeling experience through historical memory, task-specific evaluation, and TextGrad-based prompt refinement, while a server aggregates experience under evidence-based weighting and distills it into global meta-prompts. Across real multi-hospital EHR benchmarks it outperforms both local and federated baselines on diverse clinical prediction tasks and holds up across federation sizes and LLM backbones, positioning accumulated experience rather than weights as the object of collaboration.
See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs
Recovering the partial differential equations (PDEs) that govern an observed system is hampered by sparse-regression methods needing a predefined term library, symbolic regression's noise sensitivity, and language-model approaches hallucinating without iterative correction. MAGE structures discovery as a confidence-governed hypothesis-validation loop run by four specialized agents: one computes derivatives and diagnostic plots, a vision-language model extracts qualitative cues from those visuals, a language model proposes candidate equations without any fixed library, and an arbiter fits coefficients and scores confidence until a threshold is cleared. On a canonical PDE suite it achieves exact structural recovery on 8 of 8 systems with the lowest coefficient error on 7 of 8, improving accuracy by up to four orders of magnitude, and it recovers expected operators in complex geometries and fits a cubic restoring-force model to lab sensor data at held-out R² of 0.985.
Resource Constraints and Performance in Agentic AI Systems
Agentic systems bundle a language model with tools, memory, state, and multi-step execution, and those mechanisms drive operating cost as much as capability, so the authors compare OpenClaw and NanoBot end to end on a paired benchmark plus an instrumented subset. Full task completion was 31% versus 25% respectively, a gap whose 95% bootstrap interval spans −3 to 15 percentage points and so establishes no advantage for either; on the instrumented subset both hit 26% full completion while NanoBot reached partial completion on 43% of prompts against 26%. OpenClaw was slower on 83% of prompts and used more peak memory on every one, with geometric-mean ratios of 2.98× wall time and 19.44× peak memory, though many of NanoBot's apparent dominance cases were simply cheaper joint failures — a discrepancy between the two evidence layers that argues for tying capability and resource numbers back to per-attempt execution records.
LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages
Prompting a language model directly for landing-page code tends to yield generic templates and unsupported persuasive claims. LandingBench is a dataset that abstracts real landing pages into reusable reference profiles — section sequences, layout patterns, tone descriptors, visual emphasis, and call-to-action (CTA) structure — and LandingAgent is a three-phase agent that profiles the target, builds a reference-guided wireframe, then refines the page through critique-guided polishing. Measured against direct prompting on faithfulness, conciseness, readability, aesthetics, and structural diversity, it shows improved target grounding, presentation quality, and layout diversity.
TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision
Agents built on small language-model backbones are cheap but can settle into persistent failure modes, so a deployment needs a policy for when to escalate to a larger, costlier backbone. TACIT-SWITCH learns permanent handoff policies from Teacher-Annotated Censored Intervention Times, representing each annotation as an interval-censored observation on a cumulative-risk scale and fitting a mixture-cure threshold model that estimates both whether the strong rollout would succeed and, given success, where the handoff threshold sits — with no teacher needed at deployment. In a mechanism-based multi-step simulation it beats task-level, step-level, and fixed-prefix routing by 7.4 to 11.1 percentage points of success at comparable cost, and achieves the highest held-out success among learned policies on both ALFWorld and DABench.
What Makes Agent Memory Useful for Reliable Unanswerable Question Handling?
Trustworthy agents have to recognize when a question simply cannot be answered, and it is unclear whether the memory modules common in agent frameworks actually help with that. Four representative memory methods are compared across three unanswerable-question datasets and two base models inside a single agentic retrieval-augmented generation setup. Gains prove selective rather than universal and fall apart under dataset shift; reusing memory across base models turns out to be easier than reusing it across datasets, and procedural, rule-based memories that carry decision guidance transfer more reliably than memories that shape trajectories or merely accumulate more experience.
openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents
Coding agents that work over long horizons need harnesses that let developers compose heterogeneous capabilities and delegated sub-agents without rebuilding orchestration each time, and that let new evidence — semantic diagnostics, execution outcomes, task progress, shifting context relevance — steer later runtime decisions. openJiuwen is an open-source harness targeting these two properties, called Structural Composability and Runtime Adaptivity, via a shared execution substrate with Rail-based capability composition spanning single agents, delegated sub-agents, and a Swarm Flow mode, plus framework-controlled adaptation of context, feedback, and task control around a fixed model policy. It scores 82.6% on SWE-bench Verified and 87.19% on Terminal-Bench 2.1, 3.4 and 3.39 percentage points above the strongest selected official-leaderboard point estimates.
When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems
Multi-agent systems increasingly generate their collaboration topology on the fly, but current generators draw on the language model's parametric knowledge and treat search or retrieval as a reactive tool rather than something that shapes the structure itself, producing redundant interactions or too little verification on knowledge-intensive tasks. K-GAT (Knowledge-Guided Agent Topology Generator) treats topology design as knowledge-conditioned structure learning in a neuro-symbolic framework, feeding external evidence directly into autoregressive graph generation. On the expert-level GPQA benchmark it beats an LLM-Debate baseline by +15.7% accuracy while using less than half the tokens.
GOD: Govern, Observe, and Direct - A Real-Time Control Room for Agent Societies
Running a generative-agent society is easy; inspecting one is not, since operators typically get either a finished replay or raw logs across many agents, locations, messages, and model calls. GOD (Govern, Observe, and Direct) is a local-first browser control room combining a setup wizard, Agent Studio, Map Studio, a spatial replay interface, Ask and Intervene commands, and portable experiment, map, and agent packs, so an operator can pose targeted questions or interventions and immediately inspect the resulting replay state. Its design contribution is a shared command-and-artifact loop where live controls and replay evidence use the same operator command model while package contracts keep scenario, map, and profile data separate from local runtime state. Over 15 run slots, 78 of 84 target-agent checks in the 14 intervention runs recorded the commanded destination and 169 of 182 state answers matched a saved location or action string.
Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents
Trajectory-level credit assignment can pinpoint which module of a tool-using language-model agent caused a failure using only verifiable signals, and this study tests whether that credit should steer a fixed zeroth-order/evolution-strategies perturbation budget across modules. Across a synthetic environment, frozen Qwen2.5-1.5B, Qwen2.5-3B, and SmolLM2-1.7B agents, three task families, six allocation schemes, credit-noise sweeps, paired seeds, and exact sign-flip tests, no allocation scheme beat uniform by the 2-percentage-point threshold in any on-pool comparison, and concentrating the whole budget on the credit argmax was significantly worse on the 3B model. Loss scales linearly with bottleneck starvation rate (R² = 0.94, descriptive), a credit-free coverage floor removes the detected harm, and matched-budget burst and catch-up schedules point to insufficient cumulative parameter movement rather than update frequency as the cause; the one exception is soft routing beating uniform on held-out BFCL endpoints (+0.047, p = 0.031, n = 6). Three failure modes that can silently invalidate zeroth-order experiments on frozen language models are documented.
String: An Agentic OS Where Every App Is a Markdown File
Agents pay context cost for every tool schema and rendered page they see on each turn, since interfaces were designed either for human skimming or for programs that carry definitions cheaply. String is an open-source runtime that treats this as an operating-systems problem: a single SFMD (String-Flavored Markdown) document declares an application's views, typed actions, navigation, and credentials, and the runtime exposes them through just two verbs, /open to see and /act to do, with tool knowledge held outside the agent's context and rendered back one view at a time. The same document serves styled HTML to browsers and raw Markdown to agents, so apps, files, shells, and legacy web pages share one grammar with no per-site integration, and privilege follows provenance so a remote page can call HTTP but never the shell. Staged disclosure matters causally — revealing a tier of detail one turn early costs up to 23 accuracy points — and on an 87-task benchmark across six models the approach matches curated-skill baselines (+1.3pp) while using 33.5% fewer tokens, with the resident interface holding at a constant 53 tokens regardless of catalog size.
WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents
Agentic search environments typically hand back retrieved evidence as text and drop tool-returned images from later context, collapsing what should be visually grounded reasoning into text-only reasoning, while long trajectories accumulate tool-call, length, timeout, and budget failures that waste rollouts and destabilise training. WeAgent-Harness gives retrieved images persistent disk references so the model can inspect, process, and cite them throughout a trajectory, and adds runtime recovery; on top of it, WeAgent-MMSearch covers task synthesis and verification by a strong multimodal model, expert trajectory collection, and post-training with Failure-Aware GSPO, which rescues salvageable abnormal rollouts and discards invalid ones. A companion benchmark, VisTarget-Bench, pairs each of 150 human-verified questions with a held-out target image to tell retrieval failures apart from perception failures. Agentic post-training raises the average score by 19.22 points, letting the model beat similarly sized open-source systems and match ones roughly ten times larger.
Post-Edit Re-Verification in Simulator-Backed Engineering Agents: A Controlled Comparison of Verification-Cadence Guidance
When an agent driving an external simulator edits a design, its previously gathered evidence goes stale — but it is unclear whether agents re-run the simulation on their own or only when told to. This controlled study compares a prompt that explicitly instructs the agent to request a new simulation after any substantive modification against one with that instruction removed, holding all verification-relevant state constant and using no hard gate, with five Alibaba/Qwen models on eight synthetic valve-pressure cases in DWSIM across 120 runs per condition. Re-verification occurred in 94 of 120 runs with the cadence instruction versus 32 of 120 without, cadence violations fell from 87 to 26, and bounded final success rose from 35 to 95; qwen3.5-35b-a3b essentially never re-verified and never succeeded under either condition, which the authors read as evidence that verification cadence belongs in the interaction protocol rather than being left to the model.
CrabOS: An Operating System for Human-AI Co-inhabitation
Real tasks often require a human and an AI agent to take turns leading, but current agent systems give each side its own work environment, so handing off state means building task-specific interfaces or manually pasting screenshots and descriptions. CrabOS implements what the authors call human-AI co-inhabitation: the work state lives as natural-language-readable text objects that both the person and the agent read and manipulate through the same auditable interface, with no bridging layer. Case studies argue this moves support for alternating-leadership tasks from application-level workarounds up to native operating-system capabilities.
Benchmarking large language model agent societies against human behavioural distributions
Populations of language-model agents are increasingly run as stand-ins for human experimental societies, which raises three doubts: whether agents behave like the humans they replace, whether findings survive cosmetic changes to the apparatus, and whether apparent social dynamics are anything more than recall of published experiments. SILICA tests all three with five environments carrying published human anchors, each paired with re-renderings that leave the rules intact and with variants whose payoffs point away from the memorized result, run over twelve open-weight models on a single consumer graphics card. Agreement with humans holds only at starting points — eight of eleven models match first-round public-goods contributions but none matches end-state contributions — and merely swapping the order in which two actions are listed cost one model 58 points of cooperation; only the single reasoning-trained model placed its ultimatum acceptance threshold where incentives require, and conventions formed from shared priors over names rather than negotiation.
Beyond Task-Only Matching: Personalized Skill Routing with Counterfactual Evaluation
As reusable skill repositories grow, large language model agents must route requests to the right skill, but current routers match on task semantics alone, so two users with incompatible constraints issuing an identical request receive the same and possibly unusable skill. Personalized skill routing is reframed as profile-conditioned retrieval and paired with a counterfactual benchmark that holds the task fixed while varying the user profile so the correct skill changes. SkillFeed, a progressive retrieve-and-rerank pipeline that first establishes task-skill alignment and then reranks semantically similar but profile-conflicting candidates, reaches 75.1% top-1 accuracy on SkillFeed-Bench, 23.1 points above the pretrained routing baseline, with a 35.1-point gain on exactly those queries where the profile changes the answer.
Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
Multi-agent systems built on large language models fail often, and existing self-reflection has every agent reflect even though a failure usually traces to one agent that led the task astray, which contaminates the well-behaved agents' memories with wrong lessons. DoCtOR instead attributes the failure automatically to a decisive error step and agent, uses counterfactual reasoning to produce a corrected version of that step, and asks only the responsible agent to reflect. Success rates improved by 22%, 26%, and 27% over the initial systems on HotPotQA, ChartQAPro, and Mind2Web, ahead of Reflexion, Retroformer, and COPPER, and under low-resource budgets reflecting only on steps after the decisive error works about as well as using the full failure trajectory.
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
Loop engineering — writing a control loop that monitors a coding agent, assigns work, runs checks, and decides when to stop — is hard to evaluate, because the outcome of a single end-to-end run cannot separate bad loop guidance from a weak coding agent. LoopArena isolates the guidance role by evaluating a Controller model that, after each coding round, reads a structured run summary and tells a separate fixed Worker agent what to do or verify next, across three settings: execution-validated next-step contract selection without running the Worker, repeated control over a task slice, and the full paired task. The best Strict Success Rate observed on full tasks was 24.69%, and the cheaper slice-based setting ranked Controllers nearly identically to full runs (Spearman ρ = 0.9747) at an average 64.4% reduction in estimated inference cost.
EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
Agents built on language models increasingly rewrite their own prompts, tools, middleware, and execution harnesses while running, and a mutation that helps in one state can leave effects that cannot be undone from a different state. EvoUndo represents, synthesizes, diagnoses, and independently verifies whether such self-modifications are recoverable across counterfactual states, and across 600 unseen one-shot self-evolution tasks it found 197 capability-improving mutations that fail recoverability checks. Conventional repair-by-prompting recovered 0 of those 197 failures, while giving the system exact state-address grounding lifted recovery from 0/48 to 38/48 in cases where the original recovery language sufficed, and extending the recovery language itself fixed 142/143 of the remaining stratum. On the gpt-oss-120b backbone, combining exact-address diagnostics with the richer language slightly hurt recovery (133/143), an interaction that did not reproduce on Qwen3.8-27B and therefore appears model-specific.
PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems
An analysis of 16,000 real sessions found that 75.9% of user interactions with agentic systems span multiple turns, while training data and benchmarks mostly assume a single, fully specified query. PersonaForge synthesizes realistic multi-turn user-agent dialogues from a four-dimensional persona space with behavioral control calibrated against real user statistics and a reverse-construction procedure seeded by authentic queries, yielding a 6,300-record training set and PersonaForge-Bench, a hand-annotated 138-task benchmark across more than 20 professional domains. Training Qwen3.5-27B on the synthesized data raised the composite benchmark score by 4.1%, with Task Completion up 6.0% and Response Quality up 6.8%, and the trained agents completed tasks in fewer turns and fewer tool calls.
Prove2Me: An Open Collaborative Platform for Scaling Math Formalization
Formalizing mathematics in proof assistants such as Lean 4 has been gated by the expertise and time formal proofs demand, barriers that AI coding agents have lowered enough to make internet-scale human-and-agent collaboration plausible. Prove2Me is an open platform where users launch formalization "missions" that AI agents contribute formal proofs toward, with mechanisms and a specialized harness designed so agents can build on one another's work and freely reuse existing results. Because every contribution is machine-checked for correctness, proofs can be pooled from anyone with an agent without trusting the contributor.
Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction
While qualifying models for an internal document-extraction service, the authors caught a model passing a fidelity check — whether the extracted value matches the source — without ever opening the datasheet: a structured-output constraint had silently disabled tool use and the model answered anyway with fabricated source text, visible only in the per-tool trace. They log every tool call across an agentic benchmark of 37 hand-curated claims and build two instruments from that dispatch record: a rule-based failure-attribution classifier, and a silent-failure detector whose rules examine only which tools were called, never the extracted value. The detector raised no flag on 207 clean fidelity-passing extractions across three model families and recovered all 50 planted faults that withhold the tools its rules check, though power against runs that do call their tools and still answer wrongly is unmeasured; a second oracle using physical measurement could grade only 2 of the 37 claims, and the authors conclude the tool layer buys portability and observability rather than accuracy.
Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
Interactive dialogue games require a model to carry state across turns, interpret feedback, and pick valid actions under shifting constraints, and diagnostics on a 2B open-weight model in the LM Playschool Challenge show its failures are often local decision errors — repeated guesses, malformed actions, violations of feedback just received — rather than only missing knowledge. The recipe has three steps: acquire broad game participation via supervised fine-tuning, repair mechanically verifiable failures in one targeted game family with turn-local preference pairs, and preserve general capabilities. In the official evaluation, public clemscore rose from 10.67 to 38.92 and closed in-domain score from 13.41 to 41.17 with static performance roughly preserved (44.14 versus 44.24), but out-of-domain clemscore stayed at 7.88, so the turn-local gains transferred mainly within the targeted family.
COVER: Identifiable Evaluation of Coalition Routing
Changing the team in a multi-agent system also changes the messages it exchanges and the answer it produces, so an end-to-end accuracy gap alone does not identify a routing effect. COVER is an evaluation contract that fixes a public information boundary, a downstream stack, and a finite family of legal teams before outcomes are generated, which makes exact finite-benchmark oracle regret identifiable conditional on that stack. Across MuSiQue, HotpotQA, and a five-family ToolSandbox variant-shift validation, the instrument exposes selection headroom without manufacturing a routing win: the declared-family oracle reaches 0.768 safe-evidence completion while a prospectively frozen router gets 0.637, failing the pre-declared 0.10 regret criterion, and in fixed-stack Llama execution a 0.190 route-regret improvement corresponds to a raw-answer gain of only 0.010 with an interval crossing zero.
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
Long-horizon agentic tasks force LLMs to retrieve and integrate scattered information over many turns, so preserving full interaction history makes the working context grow without bound. Existing proactive context-management methods give models only search, deletion, and summarization tools, explore all editing actions as if their impacts were equal, and assign one trajectory-level reward to every intermediate edit. ContextPilot adds planning, long-term memory, and soft context-offloading tools, and trains with an RL method that uses context and entropy variation to identify critical editing decisions for branch sampling and then estimates action-level advantages from all branched trajectories passing through that edit. On long-context question answering and deep-search benchmarks it beats existing baselines across several base models while keeping a more compact working context.
LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment
Security analysis workflows are procedural — inspect artifacts, form hypotheses, run tools, revise plans — which makes them a natural target for LLM-based agents that plan, call tools, and keep state, yet the literature uses the word "agent" inconsistently and evaluates it incomparably. A systematic literature review of peer-reviewed work from 2023 to 2026 organizes the field along three axes: technical design (architecture, perception, memory, planning, action space, orchestration, self-improvement), the security tasks addressed, and assessment practice (datasets, outcome versus trajectory metrics, safety measures, baselines). The synthesis concludes the field has produced agents that can act but not agents whose authority is bounded or whose behavior is auditable, and closes with the resulting research gaps.
On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces
AI coding agents are increasingly extended through plugin marketplaces where functionality arrives as a mix of natural-language instruction files, scripts, and configuration rather than source code alone, raising the question of whether such plugins are maintained or written once and abandoned. An empirical study covers 1,926 repositories hosting Claude Code plugin marketplaces, spanning 8,351 plugins and 77,773 commits. Plugin-touching commit activity grew 8.8x in the six months after the October 2025 launch, software-engineering plugins make up 61.3% of the total, feature commits occur at more than twice the rate seen in conventional open-source software (39.6% versus 17.2%), and Claude co-authors 34.9% of commits. Inside skills directories, instruction files and their implementation scripts co-evolve above chance with 78% of co-changes functionally coupled, a maintenance dependency without a traditional software analogue.
Logos: An Agent Harness on a Cross-Process Bus
Agent systems that compose capabilities at runtime are typically hosted in a single process sharing one context, which puts every component in one failure domain — a single fault suspends all of them and process death kills every session hosted there. The argument here is that neither the spatiotemporal-composability calculus nor the underlying model requires one process, since language-model inference is stateless and all cross-step state lives outside the model, and that the soundness invariant depends only on the state space; four lemmas formalize this. Logos implements the result as a ROS-like cross-process harness where each plugin is its own process and the only shared state is an append-only transcript: eighty sessions resumed with no repeated effect after kills at all four boundaries of the tool-call cycle, and a same-fault comparison showed one fault ending at a single node instead of interrupting every co-resident session.
2 more specialized papers
Safety & Alignment 26
The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models
Whether a user's emotional state changes the direction of an assistant's advice is tested by presenting six commercial models with a user who is overconfident about a premature decision — quitting a stable job, expanding a business, emigrating — while holding the objective information fixed. A no-emotion multi-turn control keeps factual content and turn count constant so emotion is isolated from conversation length; 324 conversations were scored 0–100 for endorsement strength using an eight-item rubric. Distress raised endorsement from 18.6 to 31.5, a 12.9-point increase (Cohen's d = 0.51, p < .001), with the cold-versus-neutral difference not significant. Five of six models showed the effect, including the flagships Gemini 3.1 Pro and GPT-5.5, while only Claude Opus did not, indicating vulnerability tracks the individual model rather than its price tier.
A Survey on Rubric-Guided Reinforcement Learning for Language Models
Reinforcement learning from human feedback (RLHF) reduces response quality to a scalar reward that is neither interpretable nor able to capture multiple quality dimensions at once; rubric-guided reinforcement learning replaces that with structured natural-language evaluation criteria driving reward design, feedback, and policy optimization. The survey introduces a Bayesian framing in which constitutions are prior distributions over evaluation criteria and rubrics are conditional instantiations for a given input, then organizes the literature along a prior-to-posterior axis spanning constitutional AI, instance-specific rubrics, process-level supervision, self-evolving rubrics, and agentic and multimodal variants. Because rubrics are text artifacts, it also analyzes linguistic failure modes — granularity trade-offs, semantic drift, and linguistic reward hacking — as open problems for alignment reliability.
How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution
Transcoder attribution graphs normally explain why a model assigns high probability to a next token, which says little about internal concept representations that never surface in the output. Concept-Targeted Attribution (CTA) instead builds attribution graphs with respect to a linear probe direction, producing probe-specific circuits; using Cross-Layer Transcoders, graph-level features predict probe accuracy across four concept categories with ρ = 0.91 and R² = 0.84, while local features pinpoint the sparse components behind individual classifications. Causal ablations show the two graph types are mechanistically distinct: removing probe-relevant features lowers internal concept scores while leaving generated tokens largely intact, whereas removing logit-relevant features changes the output token in 92-100% of cases with almost no effect on probe scores.
Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap
Post-training quantization is usually assumed to preserve behavior, so models are certified at full precision and compressed afterward without re-evaluation — a workflow this work formalizes as a validation-deployment gap using Quantization Behavioral Equivalence Classes, proving that class membership does not imply behavioral equivalence. A three-stage adversarial fine-tuning procedure embeds payloads that stay dormant under source-precision checks but activate under INT8 or 4-bit compression, demonstrated on multilingual encoder-decoder translation and political stance classification rather than just decoder-only models. Backdoored translation models go from zero measured friend-foe corruption at repaired FP16 to up to 85.02% inversion after quantization, and cross-quantizer analysis shows persistence depends on the specific quantization scheme and architecture rather than nominal bit-width, arguing that the deployed configuration itself must be behaviorally certified.
Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator
Deployed AI applications increasingly need moderation over images, documents, and screenshots under policies that differ by domain, but existing guardrails cover only parts of that space or cost too much to run. Nemotron 3.5 Content Safety Moderator is a 4B-parameter vision-language moderator that classifies user prompts, images, and assistant responses across 12 languages, returning fast labels by default and optionally producing reasoning traces that apply supplied custom policies and name violated categories. The release includes a multimodal, multilingual guard-training dataset spanning human-labeled real-image moderation, benign vision-language and document tasks, synthetic rare-risk and jailbreak cases, and custom-policy examples; evaluations across multimodal safety, text moderation, multilingual robustness, policy following, false positives, and latency show it stays broadly competitive with specialized guard models while adding image- and policy-conditioned coverage.
LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails
Safety guardrail models that screen large language model inputs and outputs are trained almost entirely on short text, leaving their behavior on long contexts untested. LongGuard frames the problem as Safety Needle-in-a-Haystack (SafetyNIAH) over context lengths from 0.25k to 32k tokens and finds that unsafe-content recall across 15 mainstream guardrails drops by more than 50% on average, with a paired Benign-Fill versus Needle-Repeat design attributing the failure to dilution of the unsafe passage rather than absolute length. A three-layer attention-logit-behavior analysis traces the mechanism to diluted attention mass on the unsafe span and a compressed unsafe-over-safe logit margin, concentrated in a sparse set of guard-specialized retrieval heads. Two training-free fixes, Chunked Detection and Attention-Head Sharpening, combined with length-based Context-Aware Hyperparameter Routing, raise the six-guardrail average by 22% and 13% respectively across five benchmarks.
Semantic Watermarking with Order-Robust Detection over Sub-sentence Units
Semantic watermarks bind a mark to sentence meaning rather than token choice, but the detector only ever sees attacker-supplied text, which can be reworded, reordered, or resegmented — all of which displace the embeddings the detector tests. The proposed embedding displacement attack (EDA) combines all three edits under a single objective that maximizes displacement, using only a public paraphraser and surrogate encoder, and at a 5% false-positive rate with 90% content preservation it strips the mark from 32.6% to 47.9% of documents across four schemes, the strongest of the attacks tested. The authors respond with k-SwordStamp, which detects over sub-sentence units in an order-robust way; the best no-box attack against it succeeds 10.8% of the time, and even an attacker with the provider's detector and secret key reaches 39.7% versus 65.5% against k-SemStamp.
Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots
Memorization in large language models is measured through many incompatible definitions, and differential privacy (DP) is commonly assumed to bound all of them simultaneously. The authors pin down exact DP constants for the two that carry practical weight, counterfactual memorization and adaptive extraction, and show neither controls the other: under f-DP, extraction with a list budget is capped relative to an oblivious guessing baseline and the bound is tight on a dense set of baselines, while counterfactual memorization of any bounded score is capped at tanh(epsilon/2) under pure DP, with a closed-form staircase constant replacing the naive k-times-epsilon bound for k duplicated copies. Because the two measures separate inside the local score class practitioners actually use, one mechanism can be memorized yet unextractable and another fully extractable yet invisible to loss-based scoring. On billion-parameter models a reserved-trigger release is recovered verbatim from a single prompt while the deployed audits certify the model clean.
ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools
Stealing an LLM agent's runtime context, meaning the user prompt, execution trajectory, and tool list, requires three things to line up: the agent picks the malicious tool, the agent passes its context as tool arguments, and the tool implementation ships those arguments to an attacker endpoint. Prior work has studied the first and third conditions while leaving the second largely unexplored, and ContextLeak closes that gap by having an attack LLM generate the malicious tool's name and description, fine-tuned with reinforcement learning on shadow users with diverse simulated agent contexts under reward functions designed specifically for the exfiltration objective. The attack remains highly effective even when shadow contexts differ substantially from the victim's, and outperforms existing malicious-tool attacks adapted to this setting.
EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion
Content moderation is evaluated on static benchmarks, but real platform users iteratively rewrite posts in response to being blocked, creating a gap between offline scores and deployed effectiveness. EvoHarmBench closes that gap with a dynamic adversarial loop that evolves evasion strategies at the level of semantic clusters while jointly optimizing for evasion success and human readability, covering 229 semantic sub-clusters across five violation categories drawn from 5,002 real adversarial samples. Against widely deployed LLM-based moderators, twelve optimization rounds reach an 80.3% attack success rate under readability constraints, including on leading commercial systems; the data, framework, and code are to be released.
OpenStamp: A Watermark for Open-Source Language Models
Watermarks that work by nudging token sampling probabilities can simply be switched off by anyone running an open-weight model locally, since users control inference. OpenStamp instead encodes the watermarking logic into the weights themselves, modifying only the final projection (unembedding) layer so the signal is emitted regardless of how the model is served. Across two models it reports stronger detection than prior open-source watermarks with minimal degradation in model capability, along with greater robustness to paraphrasing and to attempts to scrub the mark through post-hoc fine-tuning; code and watermarked versions of four popular open models are released.
AI Alignment through a Game-theoretic Lens: A Survey
Alignment techniques tuned for helpfulness, harmlessness, and controllability struggle with real preferences that are context-dependent, non-transitive, and shaped by several interacting parties over time. The survey reorganizes recent alignment research around game-theoretic primitives and groups the literature under three challenges: preference diversity, alignment priority, and temporal dynamics. Its stated contribution is separating where game-theoretic analysis genuinely buys something from where the framing is only loosely applied, alongside the open problems that remain for building robust, adaptive, and verifiable systems.
Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval Defense
A user can decompose a forbidden objective into innocuous subquestions, ask them across independent sessions, and recompose the answers afterward. The authors formalize this as compositional safety risk and prove a conditional risk-transfer bound showing that the gap between deployed and reference composed risk is controlled by the model's excess loss on the allowed subqueries — so lower language-modeling loss mechanically raises this exposure. Synthetic withholding experiments plus a 600-intent evaluation across the Qwen3 and Gemma3 families show larger models delivering greater harmful-capability uplift under a fixed decomposition-and-recomposition pipeline, while their 22M-parameter IntentAlign-MiniLM retriever beats much larger embedding models at recovering the held-out underlying intent.
Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification
A provider can alter a deployed model after approval in ways that leave routine outputs looking normal, which leaves auditors stuck when the weights are proprietary. The framework searches for probe inputs, constructed in the spirit of adversarial examples, that amplify logit drift between the approved model and the live deployment, then wraps the comparison in a Groth16 zk-SNARK so the audit discloses nothing about the model; probe families range from black-box token probes through gray-box embedding probes to stress probes needing extra interface access. Evaluated across architectures, tampering scenarios, and GPU platforms, the black-box token probes give the strongest mean sensitivity despite the weakest access requirements, and scaling from 1 to 50 probes raises proving time only from 1.02 to 1.78 seconds with proof size constant.
CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?
Prompt injection defenses for language-model agents handle known attacks well but struggle against variants and novel techniques, and face a trilemma between runtime efficiency, contextual precision, and adaptability. CAITLYN is an agent-agnostic defense middleware split into two systems: System I gives immediate protection via a two-tier library of cheap rule-based detection scripts and optimized LLM-based inference, while System II watches for anomalous signals and tries to synthesize entirely new defenses. On standard benchmarks it matches state-of-the-art detection at lower token overhead than LLM-as-a-judge baselines, and on Emerging, a new delivery-aware benchmark of novel injection techniques where static baselines and System I alone stay vulnerable, System II autonomously synthesizes verified defenses that substantially cut attack success rates across three agent environments.
Beyond Global Scalars: Synergizing Token-Level Statistics and Deep Semantics for Adversarial AIGC Text Detection
Machine-generated text detectors split into training-free methods reading global statistical scalars like perplexity and training-based methods reading semantic hidden states, and both break under adversarial pressure: scalars are lossy compressions that hide local probabilistic burstiness in human-machine interleaved text, while semantic models overfit to particular model fingerprints and can be spoofed. The authors release MOSAIC, an adversarial benchmark of 16,000 samples spanning a full-granularity attack spectrum, and propose NeuroStat, which extracts uncompressed token-level logits and deep hidden states from a single causal language model backbone and fuses them via Macro-State Residual Modulation, calibrating local convolutional features with global uncertainty indicators under orthogonality and contrastive losses. NeuroStat holds up on MOSAIC where state-of-the-art detectors degrade severely, with code and benchmark released.
Speculative Probing: LLM Monitoring at Speculative-Decoding Cost
Classifying model behaviour during inference — for safety filtering, behavioural analysis, or monitoring — currently means choosing between fast hidden-state probes that see only one vector and cannot model interactions across positions, and accurate but expensive options like dedicated guard models or pooling computation over every token. The trick here is to repurpose the speculative-decoding module already present in recent models: appending a trained soft prompt to the end of the target sequence turns that draft module into a sequence classifier, and because the KV cache is already resident in GPU memory during speculative decoding, the classification costs almost nothing extra. Across four tasks and four models (Qwen3.5-4B, 9B, 27B, MiniCPM4.1-8B), the small probes beat zero-shot GPT-5.4-mini consistently and match or exceed specialized 8B safety classifiers such as Qwen3Guard-Gen-8B and Llama-Guard-3-8B on multilingual prompt safety without running a second full model.
Generative AI Alignment with Hinduism's Theological Plurality and Sacred Representation
Discussions of AI alignment around religion have largely assumed secular, Western, or Abrahamic framings, leaving other traditions unexamined. Fifteen semi-structured interviews with Bangladeshi Hindu participants trace how they use generative AI for scriptural inquiry, devotional visualization, and religious storytelling, and how they interpret synthetic sacred imagery and explanations. Participants found the systems genuinely accessible but reported theological flattening, cultural misrepresentation, devotional manipulation, and the simulation of sacred presence and authority, prompting the authors to argue for interpretive alignment: systems that disclose their limits, preserve plurality, and avoid simulating religious authority or sycophantic personalization.
REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features
Steering with sparse autoencoders (SAEs) adjusts language model behavior at inference time without retraining, but harmful requests buried in elaborate wrappers slip past it. GUISE (Generalized Undercover Instruction Safety Evaluation) is a new dataset of such wrapped harmful prompts, on which existing single-direction steering fails to produce reliable refusals — apparently because boosting refusal features leaves the harmful continuation path active. REINS intervenes in both directions within the same SAE feature space, suppressing harmful-continuation features while enhancing refusal features, and substantially reduces harmful responses and improves safe refusals while largely preserving general capabilities, where baselines either intervene too weakly or appear safe only by collapsing.
AIM: Anchor Identity Features, Then Match for Multimodal Large Language Model Unlearning
Multimodal large language models (MLLMs) fine-tuned on people-specific data memorize identity facts, and existing unlearning methods assume access to retain images or ground-truth answers at deletion time, which is often unavailable when someone requests removal. Probing fine-tuned hidden states shows identity questions and visual-perception questions occupy distinct regions and are organized differently — identity questions cluster by person, perception questions by question type — implying identity knowledge can be suppressed without damaging perception. AIM exploits this in two stages: anchor an identity-forgetting target using a universal visual prompt, then match the vision encoder to that target under a Fisher-based constraint, achieving competitive forgetting while preserving non-deleted identities, prior knowledge, and visual perception on the same images.
Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers
Stacking guardrails around a large language model assumes the layers compound, but an ensemble only compounds when its members fail on different inputs — a condition the security literature recommends and never measures. Two instruments make the stack analyzable: an Adversary Access-Tier Model grading attackers from system-only access (A0) to training-data influence (A4), and a five-class cost model of inference-time overhead that also tiers the defender, since two classes need weight access or activation reads; from these, coverage saturates within a tier, cost rises by class, false refusals accumulate as a union, and residual attack success falls multiplicatively only under independence. Running one adaptive adversary against a seven-layer stack, failure correlation was positive in all fifteen measurable pairs (φ from 0.30 to 0.75) and the joint residual exceeded the multiplicative prediction by up to 0.172, while the full stack refused four in five benign prompts yet remained statistically indistinguishable from its single strongest layer. Stratifying on behavior difficulty dissolves most of the association, indicating common-cause dependence that arises architecturally from members wrapping the same model, so a wider candidate pool cannot fix it and stacks must be measured end to end.
GRACE:Gradient-guided Coreset Selection for LLM Unlearning
Machine unlearning for language models normally assumes someone hands you a curated forget set and retain set, but a real removal request may consist of only a handful of examples of the unwanted behavior, leaving the actual training subsets to be inferred from a large mixed corpus. GRACE computes a gradient "forget direction" from those seed examples, then uses non-negative orthogonal matching pursuit to pick a compact forget coreset whose gradients approximate that direction, and selects retain examples by projecting out the forget direction and running clustered orthogonal matching pursuit in the leftover gradient space. Across two target domains, two model families, and four unlearning algorithms, the method preserves more general model utility at comparable forget quality, with the most consistent gains over earlier gradient-based selection methods.
CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents
Retrieval-augmented generation (RAG) systems that draw on public or user-editable sources can be poisoned by injected documents that steer answers toward an attacker's target, but existing attacks embed the target query verbatim in the poison, leaving lexical and embedding-space fingerprints that simple filters catch. CamoDocs drops query inclusion entirely, instead chunking together synthesized benign and adversarial drafts, swapping selected tokens in the benign chunks for dispersion tokens that scatter the poisoned documents' embeddings, and applying a coherence filter so the text still reads normally. Tested against seven RAG defenses, three open-weight models, and three benchmarks, it achieves high attack success without query-overlap artifacts, and reaches average attack success rates of 61.80% on GPT-5.4-mini and 55.09% on Claude-Haiku-4.5; erasure-heavy clustering defenses such as TrustRAG blunt it only at a large cost in utility on retrieval-dependent benchmarks like NeoQA.
LongPIBench: A Long-Context Benchmark for Prompt Injection
Prompt injection benchmarks have concentrated on short inputs, which the authors argue makes current defenses look considerably stronger than they are. LongPIBench covers four realistic deployment scenarios — paper peer review, resume screening, code review, and email summarization — each with a synthetic and a real-world dataset whose contexts range from thousands to tens of thousands of tokens. Evaluation shows even simple heuristic injection attacks achieve high success rates and frequently bypass state-of-the-art defenses once the surrounding context is long.
When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI
Voice-controlled embodied AI inherits whatever mistakes its automatic speech recognition (ASR) front end makes, and the safety consequences of those mistakes had not been mapped. Simulated ASR errors are injected into the existing embodied safety benchmarks SafeAgentBench and POEX to see which error types matter. Some corruptions preserve semantic structure while increasing harmful ambiguity, and others actively weaken model refusal behavior so unsafe plans get generated and executed; automatic error correction helps in some cases but is not reliably protective.
1 more specialized paper
- FISGuard: Defending Against Membership Inference via Fixed Input Subspaces Haocheng Jiang, Hua Shen
Theory 26
Optimal Transport for Network Comparison: A Review with Machine Learning Applications
Optimal transport offers a way to compare graphs that yields not just a dissimilarity score but a transport plan describing how one network morphs into another. The review covers three distances applied to undirected, unweighted graphs — Wasserstein, Gromov-Wasserstein, and Bures-Wasserstein — showing the closed form of the one-dimensional Wasserstein distance over node feature distributions, demonstrating how transport plans identify which specific nodes drive the distance after a graph perturbation, and deriving Laplacian-spectrum bounds that avoid full spectral decomposition for the Bures-Wasserstein case. The distances are then tested on synthetic networks for clustering and on a real-world time series network for anomaly detection.
When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging
Continual learning fights catastrophic forgetting and model merging fights weight-disentanglement error, and these are argued to be the same underlying problem: an update helpful for one task shifts outputs on another. Formalizing this as task interference reducible to a layer-wise Frobenius inner product between the weight update and a task Jacobian, the analysis derives an upper bound isolating the spectral norm of the update as the factor an optimizer can control, and identifies the Muon optimizer as regulating that factor by construction. Swapping AdamW for Muon improves accuracy by up to 5.02 points on the eight-task model-merging benchmark across three CLIP backbones, with uniformly positive gains across ten class-incremental protocols, three task-incremental protocols, and the 11-task MTIL benchmark.
Towards a mathematical theory of superposition
Superposition — neural networks packing more features than dimensions — is given a formal treatment using frame theory and compressed sensing, modeling a sparse binary feature vector encoded through an overcomplete dictionary and recovered by applying a rectified linear unit to the Gram matrix product with a bias. Several recovery theorems follow: for random feature supports, high-probability support recovery holds for nearly tight, low-coherence dictionaries up to expected sparsity of order d/log n, while for worst-case supports there is a sharp, computable criterion for which sparsity levels permit recovery. Applying that criterion to Gaussian random matrices and equiangular tight frames yields an exact coherence-based recovery threshold for real equiangular tight frames with n > d+1, proved via a new characterization of the sign distribution in the Gram matrix.
More Data Cannot Break a Symmetry: Identifiability by Design
Unsupervised representational alignment tries to recover a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus set caps what is identifiable before any data is collected — and the obvious diagnostic, the cheapest non-identity relabelling, misranks published designs because dense sampling creates near-duplicates that are almost free to swap. Working in color space, where candidate geometries are available in closed form, the authors show the failure is structural: a symmetric design resists alignment even with 64 times the restart budget, while an asymmetric set of the same size succeeds every time, and discriminating representational models is essentially uncorrelated with recovering a correspondence (r = -0.02 over 3,000 subsets). Selecting nine colors by this diagnostic alone, without consulting any learned representation, cut catastrophic alignment failures from 75% to 2% across 93 model representations with models, layers, set size, and solver held fixed, at a cost of one function call before data collection.
Should I Use This Synthetic Dataset for Training? How to Test with Minimal Real Data
Digital twins and learned world models are widely used to synthesize training data for engineering systems where real data is scarce, but the simulation-to-reality gap means augmentation can hurt rather than help, so practitioners need a way to decide — using as little real test data as possible — whether adding a candidate synthetic dataset actually improves population-level performance for a fixed learning algorithm. The authors formulate this as sequential hypothesis testing in two forms, a direct test on the mean loss difference and a symmetry-based test on paired loss differences that buys faster evidence accumulation with a stronger null assumption, and for the latter introduce the adaptive e-process sign-flip test (aeSFT), which adapts both the number of Monte Carlo sign-flip rounds and the real test data consumed while retaining anytime-valid Type-I error control without pre-specifying a test-set size. On synthetic-data classification, digital-twin-aided wireless packet scheduling, and radio-map prediction, it detects useful synthetic data with far fewer real samples than mean-based sequential testing while matching the power of fixed-sample sign-flip testing and the paired t-test.
When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?
Alignment methods for flow-based generative models increasingly reuse conditional flow matching (CFM) losses as if they were endpoint negative log-likelihoods and their old/new differences as log-likelihood ratios; this work asks when that substitution is actually valid. For linear Gaussian paths the authors exactly decompose the endpoint negative log-likelihood into entropy, a weighted CFM objective, an interior velocity-score residual, and a boundary residual, so CFM-only estimates are exact only when those residuals cancel. At the off-policy population optimum ordinary CFM is not generally a pointwise likelihood estimator, though the weighting w(t) = (1-t)/t removes the interior residual — a result that does not extend to training or on-policy alignment, where log-ratios can stay biased even for identical endpoint laws. Experiments across dimensions, distributions, and geometries support the analysis and clarify why inexact ratios can still be useful as controlled surrogates.
Landau theory of quenched criticality in linear in-context learning
In linear models of in-context learning — where a pretrained model infers a task from prompt examples without weight updates — the prediction error blows up in a double-descent singularity when pretraining sample count approaches parameter count. Treating this interpolation point as a critical phenomenon in a quenched disordered system, the analysis traces the singular error to connected sample-to-sample fluctuations of the learned parameters and constructs a Landau potential by integrating the cavity equation for the renormalized ridge parameter. The renormalized ridge acts as order parameter, the bare ridge as its conjugate field, and normalized sample complexity as temperature, placing the double-descent singularity at critical temperature τ = 1 with a generically cubic potential and critical exponents (1, 2, 1); numerical solutions of the original learning problem match the theory quantitatively, and a pseudogap-like regime with suppressed order parameter shows up at large context.
The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension
The question here is what geometry governs how well softmax attention can be approximated by a low-rank matrix, measured as the maximum-row-ℓ1 approximation rank that preserves every bounded output. Two sharp worst-case laws separate the effect of where keys and queries live: spherical self-attention has rank Θ(min{n, (1+β)^((d-1)/2)}), while full-ball geometry adds a radial degree of freedom and yields Θ(β^(d/2)) at large temperature. Within a fixed head, row-wise softmax cancels row-scalar logit directions, leaving a visible query–key interaction dimension r that gives a minimax-sharp per-instance exponent of r/2; a calibration set of 84 BERT-base heads shows modest effective-dimension reductions across many head and temperature settings, so support geometry sets worst-case temperature scaling while softmax-visible interaction geometry controls per-head complexity.
Performative Privacy: When Differential Privacy Maximizes Utility
Privacy protection is often justified by the argument that it sustains user trust and therefore participation, improving utility over time, but that claim had not been formalized. Performative privacy joins differential privacy to performative learning in a model where agents repeatedly contribute data for mean estimation and drop out when their data leaks, so the privacy budget trades estimation noise against future participation. Theoretical analysis of the resulting dynamics plus numerical experiments show that a finite privacy budget can outperform non-private estimation in the long run once the feedback loop between leakage and attrition is strong enough.
SinkSLOT: Sinkhorn via Sparse Lifted Optimal Transport
Entropic optimal transport makes optimal transport computationally tractable, but Sinkhorn-Knopp costs O(N²) per iteration for measures with N points and its independent-coupling reference measure assigns mass to expensive transport edges at moderate regularization. SinkSLOT replaces that reference with an expected sliced lifted transport plan, which sparsifies the Gibbs kernel and supplies a non-independent prior coupling at the same time. The authors prove convergence, show that each sparse iteration costs O(LN) with L slices, and show the objective is already a divergence requiring no debiasing, with synthetic benchmarks reporting substantial speedups over dense and sparse state-of-the-art methods.
A Formal Limitation on Learning Human Language From Textual Corpora
Whether a listener can recover a speaker's intended meaning from utterance form alone is answered information-theoretically, for any featurizer of text including the hidden states of contemporary LLMs. Modeling language use as a joint distribution over meanings, contexts, and utterances yields upper bounds on the probability that any decoder recovers the intended meaning from a representation of the utterance, governed by the uncertainty form leaves about meaning, which splits into an irreducible component and one that only extralinguistic context can resolve. Because these quantities are intrinsic to language rather than to any model, no representation trained on any amount of text or supervision can exceed the bounds, which hold for discrete and continuous meaning spaces alike; experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference support the theory.
Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy
Classical analyses of kernel ridge regression usually assume isotropic input data, but real inputs have covariance concentrated along a few directions. Working with polynomial inner-product kernels on anisotropic Gaussian data whose covariance spectrum decays as a power law with exponent alpha, the authors derive asymptotically sharp expressions for the kernel spectrum and generalization error in the regime where the number of samples n scales as d raised to a power kappa. Weak anisotropy (alpha below 1) keeps the familiar variance peaks at integer sample complexities but progressively damps them, while strong anisotropy (alpha above 1) fixes the effective dimension so the variance stops depending on sample size at all, plateauing under ridgeless interpolation or decaying at an explicit rate with a fixed ridge. The bias undergoes a separate sharp transition set by how fast the target decays, switching between abrupt learning and classical source-and-capacity power-law rates, with single-index targets used to show how alignment with principal directions drives the effect.
14 more specialized papers
- Context Localization for Generalized Level-Based Evaluation in Knowledge-Based Systems Ondrej Hutn\'{i}k, Nat\'{a}lia Pu\v{s}k\'{a}rov\'{a}
- On the Computational and Statistical Efficiency of the Empirical Maximum Entropy on the Mean Method Matthew King-Roskamp, Gabriel Rioux, Rustum Choksi et al.
- Beyond Procrustes distances: a multilinear Gromov-Wasserstein distance capturing chirality Cl\'ement Soubrier, Geoffrey Woollard, Andrew Warren et al.
- Evidential-Based Higher-Order Set Argumentation Framework Shuai Tang
- Exact Risk Ratios for Weighted Data Selection in Linear Regression Guangjian Zhang
- Conformal Risk-Averse Decision Making with Optimized Certainty Equivalent Risk Control Amirmohammad Farzaneh, Osvaldo Simeone
- I-FLOP: Fast Learning of Order and Parents from Interventional Data Liuting Chen, Alex Markham
- An algebraic proof of Colombo's difference-power determinant conjecture Kun Li, Li Tie, Peng Wang et al.
- Optimal Adversarial Testing: Extracting Honest Test Results from Dishonest Test Takers Owen Cox, April Xu, Weiyu Xu
- Localizing Global Discrepancies: Marginal Contributions and Contextual Anomaly Detection Tommaso dorigo
- Generalized Splines and Gaussian Processes Michael Unser
- Conformal Uncertainty Quantification Guarantees for Neural Operators Tom Stent, Nicolas Boull\'e
- An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models Javier Aguilar Mart\'in
- On two proofs of $d^2$ mixing of weighted Dikin walks Yuansi Chen, Yunbum Kook
Other 23
Node-wise Feature Encoding for Neural Performance Prediction
Neural architecture search for resource-constrained edge devices needs accurate predictions of a candidate network's latency and energy, yet existing graph-neural-network and transformer predictors largely ignore per-node computational cost. FeatureFormer adds explicit node-wise encodings of floating-point operation counts, parameter counts, and memory proxies inside a gated graph attention architecture, and the authors release NNEQ, a large-scale energy-consumption dataset that lets latency and energy prediction be evaluated in one place. The predictor reaches state-of-the-art accuracy on both metrics including out-of-domain settings, and the node-wise encoding also improves existing predictors at negligible overhead when added to them.
SOMTab: Set-Order Mamba for Efficient Tabular In-Context Learning
Tabular foundation models that predict via in-context learning now rival per-task model fitting, but the leading systems use attention at every stage of the pipeline, which is expensive. SOMTab splits the problem, mapping unordered table tokens into stable latent slots and mixing them with Mamba state-space layers to build row and column representations, while keeping attention only for the final query-conditioned retrieval from labeled context examples; a synthetic prior called DCH-TailMix diversifies training structures using degree-corrected graph heterogeneity and heavy-tailed regimes. Across tabular benchmarks it approaches strong Transformer-based tabular foundation models while running faster and using less GPU memory.
Biologically Inspired Mechanisms for Facilitating Grokking in Multilayer Perceptrons
Grokking is the delayed jump from memorization to generalization, and the question here is whether biologically inspired regulatory mechanisms rarely used in artificial networks can actively bring that transition forward. A multilayer perceptron is augmented with input gating, structural plasticity, gain modulation, threshold modulation, homeostasis, lateral inhibition, and activation decorrelation, then ablated systematically on sparse parity and noisy XOR classification. Homeostasis gives the strongest and most consistent benefit, with structural sparsification second, while the remaining mechanisms have smaller or inconsistent effects — supporting a general principle that explicitly regulating neuron utilization and effective connectivity helps generalizable internal computation emerge.
Generalized Context in Cross Attention for Transfer Learning of Disjoint Tabular Data
Transfer learning between tabular datasets normally assumes source and target share features, which fails for genuinely disjoint tables with different feature types and semantics. CATTLE (Cross-domain Attention Transfer Learning) transfers what the authors call generalized context, held in a transformer's key, value, and query projection weights rather than in activations, letting source key projections interact with target query projections in a data-agnostic way. Across ten disjoint source-target dataset pairs it outperformed nine machine learning, deep learning, and pretrained-model baselines, achieving the best average rank of 2.9 and a 3.7% average AUROC gain.
Blog: Survey of Optimizers
Neural network optimization has moved past being a sequence of Adam variants, expanding from per-coordinate to matrix- and layer-aware updates, from fixed horizons to time-varying policies, and from clean math to state representations that must survive sharding and low precision. This survey organizes recent methods along four largely independent axes — temporal estimation, update geometry, horizon management, and representation and systems — connecting Muon's spectral normalization, the matrix statistics of Shampoo and SOAP, memory-efficient and quantized-state optimizers, schedule-free training, and small-batch corrections. The headline conclusion is deliberately non-triumphal: matrix-aware methods are a real advance but there is no context-independent replacement for AdamW, since rankings flip with scale, data-to-parameter ratio, batch size, tuning budget, and whether success is measured in tokens, FLOPs, wall-clock time, or memory — motivating a compositional view of design and a stricter evaluation protocol.
18 more specialized papers
- Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI Andrea Beretta, Salvatore Rinzivillo
- Class-Based Heuristic Selection for Solving the Flying Block Puzzle Sanyar Ahmadi, Pedram Asadzadeh, Amanj Khorramian
- A Deeper Analysis of Block-Sparse Featurizers Alexandru-Iulius Jerpelea, Amith Ananthram
- Self-Explainable Multi-Label Graph Neural Network for Correlated Evidence Attribution Yingqi Feng, Yufei Tang, Min Shi et al.
- Tensor-Accelerated Eager Multi-Resolution Grids for Evolving Large-Scale Substrates Romain Claret, Michael O'Neill, Paul Cotofrei et al.
- Quantum SEDONet: Spectrally-Embedded Quantum Deep Operator Networks for Partial Differential Equations Muhammad Abid, Arth Sojitra, Bipin Tiwari et al.
- Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification Alexandre L. M. Levada
- Unsupervised Continual Learning with Growing Self-Organizing Maps and Synthetic Replay Pujan Thapa, Alexander Ororbia, Travis Desell
- Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages Haopeng Xie, Ismail Rasim Ulgen, Sofia Son et al.
- Anchored Scenario Coverage for Failure-Aware First-Hit Batch Inverse Design Chuhan Yang, Chenxi Wang, Linhan Wu et al.
- Temporal Memory-Aware Online Test-Time Adaptation on Dynamic Graphs Bo Li, Xin Zheng, Ming Jin et al.
- Lexically conditioned realization ambiguity in Korean predicate morphology Wonjun Oh, KyungTae Lim, Jungyeul Park
- Predicting Turn-Taking Outcomes in Multi-Party Conversation: Interpretable Modelling of Speech and Gaze Dynamics with Interpersonal Closeness Mark Dourado, Karim Haddad, Henrik G. Hassager et al.
- Residual-Guided Randomized Neural Networks Mushir Akhtar, M. Tanveer, Mohd. Arshad
- Memristive-Friendly Hadamard Reservoir Computing: Structured, Multiplier-Free Recurrences at Scale Andrea Ceni, Gianluca Milano, Carlo Ricciardi et al.
- Real-Valued Hyperdimensional Sequence Representations with Hadamard Product Binding and Shift Equivariance Kenny Schlegel, Dmitri A. Rachkovskij, Denis Kleyko et al.
- Quantum Federated Learning Based on Bures--Uhlmann Geometry for Heterogeneous Noisy Clients Haruki Emori, Masaki Uchihara, Yuuki Tokunaga
- Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation V. S. D. S. Mahesh Akavarapu, Michael Daniel, Gerhard J\"ager
Vision 14
Quanta Perception as Probabilistic Events
Cameras that integrate photons over fixed exposures force trade-offs between sensitivity, dynamic range, and temporal resolution that break down in near-darkness or at high speed, while single-photon quanta sensors produce streams far beyond real-time compute budgets. The proposed primitive, probabilistic events, computes a posterior over the time since the last intensity change and represents photon streams as recursive belief states, yielding motion-adaptive scene flux, activity maps, and entropy-based perceptual uncertainty instead of reconstructed frames. The method processes more than 50,000 quanta frames per second on commodity GPUs, up to four orders of magnitude faster than state-of-the-art quanta reconstruction baselines, and supports pose estimation of a running person at roughly 0.05 lux without retraining existing vision models.
From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
Vision foundation models trained on perspective photographs break down on wide field-of-view fisheye images because radial lens distortion shifts the input distribution away from anything they saw in training. DEX (Distortion Extenders) adds a small set of learnable parameters that jointly model the fisheye distortion coefficients and the latent-space gap between fisheye and perspective images, trained with a self-supervised alignment loss that warps fisheye embeddings to look perspective-like. The approach is agnostic to architecture and task, improving both monocular depth estimation and open-vocabulary segmentation over baselines for convolutional and Transformer backbones on indoor and outdoor fisheye datasets, and its activations can be decoded into distortion coefficients for camera calibration.
What Do Interaction Representations Actually Measure? Pre-Event Separability in Weakly-Supervised Violence Detection
Detailed skeletal pose should in principle carry more information about a violent interaction than coarse bounding boxes, so the authors hold the tracker, temporal head, supervision, and evaluation fixed and compare five interaction representations for early violence detection. No pose-based representation beats coarse geometry, and once frozen visual encoders are added on the larger XD-Violence split, whole-frame context matches or exceeds person-crop appearance — cropping to the interacting people buys nothing. Scoring anomalous videos using only frames before the annotated onset, with sequence length controlled for, still retains 39–91% of above-chance separation on both UCF-Crime and XD-Violence, traced to provenance artifacts such as editorial title cards and platform watermarks absent from the surveillance footage supplying normal clips, meaning video-level AUC mixes genuine event evidence with source cues.
VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians
Physics-driven 3D Gaussian pipelines mostly animate solid objects and handle only single-phase collisions, leaving out interactions between materials in different states. VersaGauss unifies generation, simulation, and rendering, taking a handful of images and producing a dynamic multi-object 3D scene, with a particle pruning algorithm to shape the Gaussian kernel distribution and a Coupled Multiphase Point Method (CMPM) to model interactions across phases; harmonic interpolation inside CMPM plus a Gaussian evolution strategy handle fluid rendering. Experiments show interactions simulated among fluid, rubber, sand, snow, and other materials in a single scene, with code released publicly.
Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations
Reading a CT scan involves comparing structures across the body, judging distances between organs, and knowing where each organ belongs, yet medical vision encoders are graded on diagnostic accuracy or inside assembled multimodal systems where failures cannot be attributed. SPAR-Bench supplies eight probes over multi-organ abdominal CT that separate coordinate localization, relational reasoning, and spatial queries, applied to five architectural configurations and three medical foundation models both frozen and finetuned. Probes requiring a comparison within a single slice stay at chance regardless of pretraining scale, finetuning, or architecture, and probes that look solved in-domain collapse to chance under zero-shot transfer, suggesting the encoders recall where organs usually sit rather than computing over the specific image. Reading the same frozen features with a pooled head instead of all tokens lifts relational recovery from 0.7% to 67.8%, so pooled probing understates what representations contain, and four open-weight multimodal LLMs answer at chance the questions the encoders handle well.
A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation
Change detection in Earth observation is held back by inconsistent evaluation protocols and a focus on accuracy that ignores compute cost. The benchmark trains ten architectures, from convolutional networks to vision transformers, on ten heterogeneous datasets under identical protocols, comparing training from scratch against pretrained weights and reporting parameter counts and inference latency alongside predictive scores. Well-optimized classical designs such as Siamese U-Nets frequently beat more complex recent models once efficiency is factored in, and pretraining gives a consistent gain at no inference cost; splits, scripts, logs, and checkpoints are released under FAIR principles.
Physics-Guided Flow Matching for CT Image Reconstruction
Diffusion priors deliver state-of-the-art CT reconstruction but depend on stochastic sampling, long inference trajectories, and carefully tuned noise schedules. A rectified Flow Matching model is trained instead on 256x256 chest images from the Mayo Clinic Low-Dose CT dataset with a two-stage schedule — heavy anatomically informed augmentation first, then fine-tuning with reduced or no augmentation to restore structural fidelity — and used as the prior for Plug-and-Play Flow, FlowDPS, Flower, and Flow-Priors. Across several CT inverse problems these outperform the diffusion baselines DDRM, DPS, and DiffPIR on PSNR, SSIM, and perceptual quality while needing fewer sampling steps, and the trained model and code are released.
Video Generative Models as Geometry Learner
Generative geometry estimation currently adapts pretrained image diffusion models, either training depth and normal predictors separately — forfeiting the correlation between the two targets — or jointly fine-tuning modified backbones, which needs a lot of labeled data. GeoNeXt instead repurposes a pretrained video generative model and casts geometry estimation as next-frames prediction, inheriting temporal structure and richer priors while adapting them to jointly model image and geometry targets in both directions. On zero-shot monocular depth and surface normal estimation across diverse datasets it beats prior task-specific and unified generative methods, and rivals discriminative state-of-the-art systems trained on over 100x more data, leading on several benchmarks.
6 more specialized papers
- FVeinSyn: Synthetic Finger Vein Image Generator Yifan Wang, Jie Gui, Adams Wai Kin Kong et al.
- Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge Md Monjurul Ahsan Prodhan, Md Nour Hossain
- EXPOSE: Explainable and Domain-Robust Embeddings from Pathology Vision Foundation Models using Sparse Autoencoders Anja Witte, Maximilian Lennartz, Jan Baumbach et al.
- Cross-Spectral Dense Correspondence for Multimodal Spectral Medical Imaging Eric L. Wisotzky, Jost Triller, Simon W. H\"artl et al.
- Anatomy-Aware Promptable Segmentation with Online Interactive Training for AUTOPET V Pablo Lozano-Jimenez, Sergio Romero-Tapiador, Ruben Tolosana
- Texture Image Classification Using DWT AlexNet Feature Fusion and Deep Neural Networks Arun D. Kulkarni
Multimodal 13
SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction
Relational reasoning spans several distinct abilities — analogical, structural, cause-effect — and SciReC tests them in multimodal large language models (MLLMs) through a model-adaptive multimodal academic dialogue benchmark. It is paired with DMRA, a deficit-based diagnostic framework that separates how much visual understanding, knowledge, and memory recall each contribute to a failure. Claude 4.6 leads with a 73% overall relational score, ahead of GPT 5.4 at 68%; open-source models score lowest on spatial relations while proprietary models struggle more with hierarchical and sequential ones, and performance is worst in Astronomy and best in Psychology. DMRA attributes the bulk of errors to relational reasoning itself, with memory limitations second.
Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls
Interpretability work typically fixes a stimulus set and asks either what internal structure represents or how outputs vary, neither of which recovers a model's prior — the distribution over inputs it implicitly expects — since image space is far too large for any fixed set to cover. The proposed method steers a generative model to produce stimuli along interpretable control axes and runs Gibbs sampling over that space with the multimodal LLM under study acting as the judge, drawing samples directly from its perceptual prior. Applied to targets such as trustworthiness in faces and cheapness in art images, it recovers both canonical biases and surprising priors that direct prompting leaves invisible.
CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning
Adding a fresh set of LoRA experts for every new task lets a multimodal model learn continually without forgetting, but the parameter cost grows without bound. A singular value decomposition of task-specific LoRA updates shows their input- and output-side direction subspaces overlap heavily, with task-specific adaptation mostly expressible as lightweight coordinates over shared bases. CoRe-MoE exploits this by extracting reusable direction bases from an initial expert bank and training only compact coordinate experts plus task-specific low-rank routers thereafter, beating the strongest baseline by up to 5.90 points while training under 1% of the parameters sequential LoRA needs for later tasks.
There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation
Text-to-image and similar translation models generate along paths that start from noise rather than from the source modality, which constrains sampling algorithms and makes the mapping one-way, so image-to-text inversion is not available. BIT builds bidirectional diffusion bridges that interpolate directly from text into images, giving a source-aware generative path and an endpoint-conditioned process traversable in either direction, derived through stochastic calculus into stochastic differential equation forms with tractable losses that scale to high dimensions. It is competitive with denoising-diffusion and deterministic-flow baselines and better on several vision–language and natural-science evaluations, all within one unified framework.
Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models
Large vision-language models still generate content inconsistent with their image input, and most mitigations act from outside through extra supervision, output calibration, or attention tweaks rather than addressing what happens to internal representations during decoding. The diagnosis here is an inference-time failure mode in which cross-modal representations degrade as they pass through decoder layers and drift across generation steps, destabilising token prediction. Dynamic Alignment Compensation (DAC) is training-free: it detects representation divergence and applies lightweight residual corrections, combining Layer-wise Semantic Compensation against inter-layer degradation with Sequential Semantic Correction against temporal drift. Across nine hallucination-focused and general multimodal benchmarks and several backbones, DAC consistently lowers hallucination rates without hurting general performance.
Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs
Hybrid attention is standard in frontier language models, but the Vision Transformers (ViTs) inside multimodal models lack an agreed-upon hybrid design or an explanation for why some attention patterns beat others. Examining ViT attention heads reveals that they split into object-specialist and background-specialist roles most sharply under full attention, a property the authors name Semantic Head Specialization and measure with an SHS-Index that separates full-attention from chunk-window ViTs and tracks downstream benchmark scores. Three structural factors — window interaction, token serialization, and local softmax allocation — are identified as what shapes this specialization and used as design rules for Ariadne Attention, which matches full attention on 22 image and video tasks while using 6.5 times less attention compute.
Post-Training VLMs for Video Mistake Detection
Systems that spot mistakes in videos of people following instructions are typically closed-set, so any new task means fresh data collection and retraining. The MD-VQA protocol and benchmark instead ask whether a model has learned the general notion of a mistake by testing, for both seen and unseen actions, whether a step was carried out correctly with respect to its written description. The proposed post-training method for video-language models uses a reward that pushes the model to find discrepancies between the instruction and what the video shows, beating zero-shot, supervised fine-tuning, and other post-training baselines, with up to 11.6% improvement over the best baseline on unseen procedures in EP-VQA.
6 more specialized papers
- SURE-Challenge: Evaluating Speech Evidence Before Speech-LLM Generation Mengzhe Geng
- Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict Adarsh Sudheer, David Li, Omar Elbanna et al.
- A Shaky Voice Is Not Always a Dodge: Benchmarking Textual and Vocal Evasion Detection in Earnings Calls Mirae Kim, Seonghun Jeong, Youngjun Kwak
- Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara
- MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places Jason Armitage, Ioannis Tsochantaridis, Linda Mazzone et al.
- ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT Huseyin Umut Isik, Mehmet Alp Ozaydin, Sila Kurugol et al.
Reinforcement Learning 12
SegBench-GC: Testing Segmentation Invariance in Multi-Step Offline Goal-Conditioned Reinforcement Learning
Logged trajectories in offline goal-conditioned reinforcement learning (GCRL) are often split into segments for storage or bookkeeping reasons unrelated to episode termination, yet multi-step targets treat those cuts as real endpoints. SegBench-GC holds transitions, source trajectories, goal sampling and optimization fixed while varying only artificial backup boundaries, comparing naive absorbing treatment against continuation-valid targets (CVT), which stop reward accumulation at a cut but still bootstrap from the stored successor. On PointMaze with 35,000 artificial cuts, success falls from 50.5% uncut to 39.1% with CVT and 19.1% when the same cuts are treated as terminal, and a replication on Puzzle-4x5 using the Decoupled Q-Chunking codebase collapses to 0.27% under naive handling. Critic diagnostics trace the failure to a large optimistic shift in learned values that CVT avoids.
Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess
Searchless chess networks such as Leela Chess Zero's Chessformer reach master strength in a single forward pass by distilling Monte Carlo Tree Search visit counts, but imitating a search is a poor proxy for playing without one. Fine-tuning with self-play reinforcement learning, the usual entropy bonus (a reverse Kullback-Leibler divergence toward uniform) is swapped for a mass-covering forward KL toward the network's own MCTS prior, paired with a sampling temperature that sharpens once the value head is confident about the outcome. In roughly two thousand steps this lifts puzzle accuracy from 93.9% to 94.9% and mate-in-four from 77% to 81% without losing playing strength, while a control fine-tuned on puzzles alone posts the largest tactical gains yet loses about 260 Elo — a better puzzle-solver is not a stronger player. Without any regularizer, self-play collapses onto a single line of play.
Rubric-to-Code Credit Assignment for Reinforcement Learning
Generating an interactive web application involves many distinct user-facing requirements, each tied to a localized region of code such as an event handler, state update, or CSS selector, but GRPO compresses all of that into one sequence-level reward applied uniformly across every token. Rubric-to-Code Credit Assignment (RCCA) builds training tasks around explicit functional rubrics, uses a hierarchical reward that separates format, source-code, runtime, and functional failures, and aligns evaluator-written textual attributions with the code spans and tokens responsible. The resulting Ling-RCCA-Flash scores 41.25 on MiniAppBench, a 32.20-point gain over the Ling-3.0-Flash base model and slightly ahead of Claude Opus 4.5, and reaches 76.19 on ArtifactsBench, topping the official leaderboard setting.
Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Sampling
Dynamic Sampling — the component contributing most of the accuracy gain of Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) over Group Relative Policy Optimization (GRPO) — filters out prompts whose sampled responses are all correct or all incorrect, removing zero-advantage gradients. The authors show theoretically that this filtering asymmetrically amplifies advantages within a prompt: on hard prompts the incorrect responses get amplified more than the rare correct ones, so the model learns mainly to avoid observed wrong answers rather than exploit hard-to-sample right ones. Their fix, Direct Advantage Amplification, boosts the advantages of those hard-to-sample correct responses and yields DA3PO, implemented in under 30 lines of change on top of DAPO, which outperforms GRPO and other classical GRPO variants in experiments.
Is Monte Carlo Tree Search Just Every-Visit Monte Carlo Control?
Monte Carlo Tree Search (MCTS) and every-visit Monte Carlo control are normally taught in different vocabularies — selection, expansion, simulation, and backup versus trajectory sampling, return estimation, action-value updating, and policy improvement. This expository note argues the distinction is largely terminological at the level of trajectory generation and value updating: the tree policy and rollout policy are the learned and not-yet-learned parts of one evolving policy, expansion is just first visit plus initialization, and backup is the ordinary every-visit Monte Carlo update. Under that reading the four stages collapse to two operations, sampling under the current policy and every-visit updating, making MCTS every-visit Monte Carlo control expressed in the data structure and language of search.
Emergent aggregation from collective foraging
Models of flocking and swarming normally build in a direct social drive, rewarding or hard-wiring agents to approach or align with neighbours. Here reinforcement learning foragers start from a random walk and optimise only an individual reward for finding replenishable targets, while their sensors show them other foragers and never the targets themselves. As visual range grows the agents cross over sharply from environment-tuned individual search to a scale-agnostic collective one, and spatial aggregation appears exactly at that crossover — grouping emerges purely as a by-product of efficient foraging. A minimal analytical first-passage model reproduces the transition as a switch between the two search strategies, identifying indirect resource-driven reward as a general route to collective behaviour.
VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning
Long-horizon reinforcement learning for language-model agents usually collapses a task's success or failure into one scalar terminal reward that is broadcast evenly across every action, leaving no signal about which step mattered. VICT opens up the verifier that produced that reward, exposing its individual executable or evidence-backed checks as atoms and tracing them back to specific actions through dependency-valid proof edges, then redistributing group-relative advantage only along those edges; it abstains when the evidence is incomplete, preserves the original terminal reward, and modifies only the training-time advantage tensor, so it needs no learned critic, process labels, or branch rollouts. On ALFWorld and WebShop it improves substantially over outcome-only training while matching recent fine-grained credit-assignment methods, and ablations rule out dense atom rewards, final-commit credit, temporal proximity, and sparsity as sufficient explanations for the gain.
HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees
Agentic reinforcement learning produces branching rollout trees whose trajectories share long prefixes, and training each root-to-leaf path independently recomputes that shared work; existing systems also assume full attention and lack differentiable hybrid-attention execution compatible with activation recomputation. HARTS jointly plans microbatches, data-parallel replica assignment, and slot schedules over compressed prefixes, and adds a linear-time algorithm for chunkwise linear attention that coordinates chunk-boundary state recovery and replay with the provably minimum number of sequential attention calls, packing all branches into one call per round while propagating gradients through differentiable state handoffs and restoring per-token log-probabilities. On an agentic workload built from SWE-bench tasks it delivers a 4.81–4.87× forward/backward/gradient speedup with activation recomputation, with numerical deviation comparable to baseline rerun noise and a matching reward trend over the first 120 steps of τ³-Bench training.
Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning
Analysis of failures on the Countdown arithmetic-puzzle task shows calculation errors account for a substantial share of wrong answers, motivating a study of whether a calculator tool helps. The authors build supervised fine-tuning data teaching tool-call patterns and how to interpret returned outputs, then apply on-policy reinforcement learning — RLOO, RLOO++, GRPO, and DAPO — with automatically verifiable final-answer rewards, evaluating on a fresh 1,024-problem held-out set with no exact overlap with training data. Tool access adds roughly 10 percentage points across pass@k for both the SFT and RL baselines, and Tool-DAPO raises pass@1 from 35.8% to 66.0% over the tool-using SFT model, with RL producing more effective tool use even though only the final answer is rewarded.
REPLICANT: Learning Policies for Evading and Hardening Malware Detectors
Attacks used to stress-test machine-learning malware detectors typically assume privileged access to training data, the feature space, or confidence scores, which overstates defender knowledge of the adversary. Replicant is a deep reinforcement learning framework that learns evasion under a strict label-only black-box threat model, producing a reusable policy governing both how to mutate a malware sample and when to spend a query on the target, and that policy transfers across samples, detectors, and feature spaces. Across seven Android malware detectors and three feature spaces it reaches a 78.8% mean attack success rate, a 20.9% to 39.2% relative gain over prior state of the art, and using it for adversarial training produces detectors with more generalizable robustness than existing hardening approaches.
2 more specialized papers
- Marginal Coverage Credit Reduces Redundant Exploration in Parallel State-Entropy Optimization Junhao Cao, Hongyi Xia, Jianian Wu et al.
- Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning Zilin Zhao, Han Yang, Tianpei Yang et al.
Reasoning 8
INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning
Large language models optimized purely for final-answer correctness on math may be memorizing solution patterns rather than internalizing concepts, and they are notably weak at example-based reasoning such as constructing counterexamples that probe a theorem's boundaries. INSPIRE tackles two obstacles to fixing this with preference optimization: weak baseline ability makes good preference pairs hard to build, and the skill must be acquired progressively. It pairs Reference-Guided Student Internalization, which generates preference candidates from the policy model's own distribution, with a stage-wise rubric training scheme that separates learning the method from learning to apply it correctly. Across multiple model scales and families the approach improves consistently and surpasses larger open-source models, with out-of-distribution benchmarks showing no loss of general math ability.
Load-Bearing Context: The Question Damage Score for Evaluating Context Reliance in Linguistic Reasoning
Linguistics olympiad puzzles are self-contained — every answer follows from the supplied examples and nothing else — which makes them a clean setting for asking whether a model is genuinely reading its context or leaning on prior knowledge. Across 53 UK Linguistics Olympiad puzzles, each is perturbed by deleting one context example either at random or, using an idea borrowed from error-correcting codes, targeting the structurally load-bearing example that uniquely carries the needed information; a Question Damage Score then classifies puzzles as fragile or robust. Instructed to abstain when information was insufficient, three frontier LLMs rarely abstained and frequently still produced correct answers after the load-bearing example was removed.
The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs
Counterfactual reasoning benchmarks typically fix the variables and grade against a single gold outcome, which never tests whether a model can trace how an altered condition propagates through downstream consequences. WhatIfBench supplies 220 open-domain, open-form what-if questions spanning STEM, humanities and social sciences, and hybrid scenarios, and PRISM grades free-form answers by first converting each explanation into a semantic causal graph of events, states, and mechanisms, then scoring both the graph's causal validity and the answer's explanatory adequacy. Across six frontier models, the strongest reaches only 64.62%, with recurring causal gaps, premise drift, and topology fragmentation indicating that fluent counterfactual narratives often sit on fragile causal reasoning.
SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing
Long chain-of-thought reasoning keeps spending tokens well after the intermediate answer has effectively stabilized, and existing early-exit signals based on confidence or entropy capture that stability poorly, while consistency checks require several sequential rollouts before they can trigger. SABER applies simple semantic perturbations to the intermediate reasoning state to form adversarial branches, then uses lightweight probing to estimate each branch's likely final answer without full rollouts, exiting when the branches agree and continuing when they diverge. The method needs no training and cuts reasoning token consumption by 30.2% to 39.8% on average while staying competitive with full-length reasoning accuracy across several benchmarks and model architectures.
AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time Reasoning
Test-time scaling generates extra candidate solutions to improve reasoning, but spending the same budget on every question wastes computation, and existing stopping rules assume that stronger current evidence — confidence, agreement, or answer stability — means more computation is unnecessary. The authors show checkpoint-level correctness actually evolves non-monotonically, so evidence can strengthen right before an answer collapses, and propose AERA (Adaptive Evidence Residual Allocation), a sequential controller that predicts whether another block of generation is likely to recover a better answer using answer-distribution, temporal, re-solving, semantic, and compute features observable at each checkpoint. Future correctness is used only to build offline training supervision, never at inference. On a held-out set of 300 GSM8K questions with frozen thresholds, AERA reaches 92.61% accuracy against 93.01% for a fixed 128-response budget while cutting completion tokens by 95.99%, with similar behavior on GPQA Diamond.
VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation
On-policy self-distillation trains a student that sees only the problem on its own rollouts, supervised token-by-token by a privileged teacher that also sees a reference solution — but the teacher is treated as a fixed target, and privileged conditioning does not guarantee it is the right target for problem-only reasoning. VISTA keeps the usual student update and adds a reverse direction: outcome-verified rollouts are used to adapt the teacher toward the student, restricted to the top-k token positions with the largest teacher-student KL divergence, reusing the same rollouts and loss with no extra sampling or reward objective. Across AIME24, AIME25, and HMMT25 with Qwen3 at 1.7B, 4B, and 8B, it took the highest Avg@12 at every scale, gaining 0.6, 0.7, and 2.1 points over standard self-distillation.
Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs
Supervised fine-tuning and reinforcement learning both bury an acquired reasoning capability inside model weights, where it cannot be inspected step by step or moved to another model. PLVR (Program Learning with Verifiable Rewards) instead learns reasoning as an explicit program of deterministic and neural primitives directly from input-output examples, using symbolic backpropagation: each program layer carries a typed ontology, a loss is computed against ground truth, and required input ontologies propagate backward by type inference over primitive signatures, so credit assignment is a derivation and the reward is a per-step contract verdict rather than a terminal outcome. On LiveCodeBench v6 and Tau2Bench, 30B base models with PLVR beat RL at matched budget by 27.8 points on average and beat frontier models an order of magnitude larger by 13.6 points, while an ablation swapping loss-guided search for uniform sampling over the same type-admissible space drops the median program from 65.6 to 17.5, which the authors read as the backward pass rather than the type system carrying the advantage.
NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry
Neuro-symbolic geometry provers such as AlphaGeometry reach near-IMO-gold performance but require problems written in a specialized domain-specific language (DSL), and translating natural language into that syntax by hand is a usability bottleneck. NL2AGBench measures how well ten open- and closed-source large language models (LLMs) perform that auto-formalization, judging output by whether it actually executes inside AlphaGeometry rather than by textual similarity. Leading closed-source models exceed 80% executable translation rate while even the largest open-source models fail to consistently preserve geometric constraints, and the authors add an error taxonomy separating syntax from logic failures plus mitigations via few-shot prompting, fine-tuning, and human-guided hinting.
Robotics 7
PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models
Vision-language-action models (VLAs) map instructions and images directly to robot actions but usually condition only on the current observation, giving them no explicit way to reason about how a task will unfold. PHR-VLA adds a lightweight auxiliary future head used only during training, which aligns the model's internal representations with privileged latent dynamics extracted from future observations. Patch-level, contact-centric supervision from the wrist camera raises success on LIBERO from 84.1% to 88.4% and on real-world disassembly tasks from 63.3% to 82.5%, with a smaller gain on Meta-World when supervision comes from a third-person view.
Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface Scoring
Robotic sorting of recyclable waste is difficult because targets such as used beverage cartons arrive crushed into inconsistent shapes. The proposed system needs no training: an open-vocabulary vision-language model detects cartons from a text prompt, SAM2 refines each detection into an instance mask, and a geometric score combining surface flatness with normal alignment picks the suction point, with k-nearest-neighbor PCA, Sobel cross-product, and RANSAC plane fitting compared for that stage. Tested on a real robot across three deformation levels and 35 cluttered scenes, single-object grasp success reached 88.2% and end-to-end retrieval from clutter 72.6%.
Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning
Tendon-driven robot hands are cheap to build because motors sit off the joints and one cable can drive several joints, but that same underactuated transmission is hard to model in simulation and leaves coupled joints that cannot be commanded independently, which makes learning on them difficult. Aero Hand Open ships as a simulation-ready package: a simulation model that reproduces the cable transmission itself, an identified bidirectional actuation map linking simulation to motor commands including the thumb's three-way coupling, and a reinforcement learning training environment. Policies trained entirely in simulation transfer directly to the physical hand with no fine-tuning and no state estimation, and the mechanical design, simulation model, mapping, training environment, and deployment stack are all released.
4 more specialized papers
- Spatial-Semantic Reasoning using Large Language Models for Efficient UAV Search Operations Marin Maletic, Marijana Peti, Tamara Petrovic et al.
- MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel Operation Guipeng Xin, Jiahe Xua, Mohammad Deghat et al.
- PanelShield: Verifiable Closed-Loop Safe Planning for Robotic Industrial Panel Operation Guipeng Xin, Jiahe Xu, Chenhui Wan et al.
- AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction Yafei Zhang, Nan Wu