Monday, September 21, 2026
Highlights
Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation
Convolutional sparse coding (CSC) suppresses redundant image components while keeping signal content, but its sparsity coefficient is normally fixed and tuned by hand. This framework unfolds the CSC optimization using the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) and makes the sparsity coefficient a differentiable variable learned jointly with the network weights, interpreted through the information bottleneck as the trade-off between compression and information retention. A label-free post-training step then adjusts the compression strength for corrupted inputs while the main network stays frozen. On CIFAR and ImageNet the method reports competitive clean accuracy and greatly improved robustness to input perturbations.
Convolutional sparse coding (CSC) layers give networks an explicit knob for discarding redundant signal content, but the sparsity coefficient λ is normally a hand-picked constant. TA-CSC treats λ as a per-layer variable learned through unrolled FISTA iterations, framed as the information bottleneck's compression-versus-sufficiency trade-off, and then re-tunes only λ on unlabeled corrupted inputs after training.
- Each CSC layer unrolls 2 FISTA iterations and backpropagates through the soft-threshold to update λ (kept positive via Softplus) jointly with the dictionary, under a loss that adds a normalized ℓ1 penalty on the sparse codes (γ = 0.001) to the task loss.
- For corrupted data, all network weights are frozen and only λ is updated on 100 unlabeled samples by minimizing a relative reconstruction error divided by λ, which pushes compression up until reconstruction fidelity degrades; the adapted λ rises monotonically with corruption severity.
- Replacing every convolution in
ResNet-18with these layers reaches 97.65% onCIFAR-10, 80.76% onCIFAR-100and 72.53% onImageNet-1K, versus 95.54%, 77.82% and 68.98% for the baseline, while swapping only the first layer still gives 96.18%, 79.63% and 71.12%. - On
CIFAR-10-CGaussian noise, accuracy goes from 44.43% (ResNet-18) to 53.98% without adaptation and 68.23% with it, beatingSDNet-18with per-sample λ tuning (64.92%); onImageNet-CGaussian noise it goes from 22.73% to 30.93%. - The all-layer variant is expensive (88.6 GB and 153 samples/s on ImageNet versus 24.1 GB and 2100 for
ResNet-18), evaluation covers onlyResNet-18classification under noise-type corruptions, and the authors concede the link between λ and the information bottleneck is empirical rather than proven.
Recursive Language Models Generalize Out of Domain
The question studied is when restricting what a language model can see improves learning, comparing standard chain-of-thought (CoT), which reads the full reasoning trace, against recursive language models that solve each subtask in an isolated context. In-distribution, CoT can efficiently simulate the recursive rule, so its generalization guarantee differs only by a constant factor and recursion offers little. Out of domain, however, CoT can fit the training data through shortcuts that depend on context outside the current subtask and break when those tokens change, a failure mode that context isolation rules out. Because simplicity bias selects the shortcut even though CoT's hypothesis class contains the correct rule, the authors argue that covering the right rule is not enough for out-of-domain reasoning, in contrast to classical learning theory.
Chain-of-thought models read their entire reasoning trace, which makes them a strictly more general learner than recursive language models that solve each subtask in an isolated context — yet this work argues that the extra generality is exactly what hurts out of domain, because it lets the learner fit training data with shortcuts that depend on context outside the current subtask.
- The comparison pits standard
CoT, which conditions on the full trace, against recursive language models (RLM), which deliberately restrict each subtask to its own isolated context and therefore cannot see tokens belonging to other subtasks. - In-distribution, recursion buys little:
CoTcan efficiently simulate the recursive rule, so its IID generalization guarantee differs from the recursive learner's by only a constant factor. - Out of domain,
CoTcan reach low training error by leaning on context outside the current subtask, a shortcut that breaks as soon as those surrounding tokens change, whereas recursive context isolation rules out this failure mode by construction. - The key theoretical point is that
CoT's hypothesis class still contains the correct recursive rule, but simplicity bias selects the shortcut over the truth — so covering the right rule is not sufficient for out-of-domain reasoning, in contrast with classical learning theory. - The abstract reports no concrete benchmarks, model scales, or accuracy numbers, so the practical size of the gap and how well the analysis transfers to real-world tasks without a clean subtask decomposition remain unclear from the summary alone.
When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
As model-written peer reviews enter training corpora, later AI reviewers may learn from earlier ones, and this study simulates one step of that loop by fine-tuning Llama 3.1 8B on official ICLR reviews from 2018-2023 and then training four successors on ICLR 2024 data with varying mixes of official and model-generated reviews. Adding synthetic reviews compresses rating distributions and reduces semantic diversity both within a paper's reviews and across the corpus, a pattern the authors name scientific-judgment collapse. To counter it they introduce TrustReviewer, an open-source reviewing system that trains in a single stage on a corpus curated to remove low-quality and degenerate supervision, and applies paired activation steering at test time to correct residual collapse without further training or expert annotation.
As LLM-written reviews leak into public corpora, future AI reviewers may end up trained on their predecessors' judgments. The authors simulate one step of that loop and find that synthetic supervision narrows both ratings and review content, which they call scientific-judgment collapse. They then propose TrustReviewer, which pairs a curated training corpus with test-time activation steering to counter it.
- A
Llama 3.1 8Breviewer fine-tuned on official ICLR 2018–2023 reviews generates reviews for ICLR 2024 papers, and four successors are then trained with 0, 1, 2, or 3 of each paper's three reviews replaced by synthetic ones (0/33/66/100%) under identical LoRA settings, with all models evaluated on 2,000 held-out papers. - Introducing just 33% synthetic reviews cuts rating standard deviation from 1.63 to 1.44 and entropy from 2.31 to 2.14 (official reviews: 1.73 and 2.38), while mean ratings move non-monotonically (5.30, then 5.85, then 5.70 at 100%), so the effect is narrowing rather than a drift toward leniency or harshness.
- Semantic diversity shrinks too: same-paper pairwise embedding distance falls monotonically from 0.159 to 0.142 (about 11%), and corpus-level spread falls from 0.609 to 0.579 (about 5%), with most of that contraction arriving at the first synthetic step.
TrustRevieweris trained in a single stage on 112,743 curated paper–review pairs (~1.9B tokens) from ICLR 2018–2025 and adds a steering vector (the mean last-token hidden-state difference between official and model-generated reviews of the same 5,000 papers, applied at the final layer with α=0.15), reaching 75.40% exact match and 1.079 MAD versus 73.10% forOpenReviewer, 61.85% forQwen3.6-35B-A3B, and 33.93% for the base Llama, with steering itself contributing +1.55 points and lifting rating entropy from 2.13 to 2.18 while slightly lowering same-paper diversity.- The collapse study covers only one recursive step, one 8B model family, and one venue, and
TrustRevieweris benchmarked against external baselines rather than tested inside the contaminated-data loop; the authors also caution that diversity is not quality, that agreement with official ratings is not correctness, and that embedding distances are only a proxy, with no human evaluation of critique validity.
SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
Agentic coding benchmarks judge patches with held-out test suites that are inherently incomplete and increasingly vulnerable to memorization, while formal verification has so far only covered standalone tasks whose specifications are handed to the model. Benchproofer turns a coding task with a known-correct patch into a formally verified one by writing a specification for the new code, summarizing the existing functions it calls with axioms, and admitting an instance only when mechanical and adversarial gates agree; applied to SWE-bench Verified it produces SWE-Proof, 500 real issues checked by proof rather than test, and extends to SWE-bench Pro. A quarter to a half of test-passing patches from two frontier models admit counterexamples, and supplying a correct formal specification lifts Opus 4.8 from 85% to 95% resolution. Writing the specification is the hard part: models asked to produce their own gain nothing over an unaided baseline, only 62% of those specifications pass audit, and the usual failure is faithfulness — constraining part of the required behavior and leaving the rest free.
Coding benchmarks judge patches with finite hidden test suites, which are incomplete and open to memorization, while formal verification has so far only covered small standalone tasks with specifications supplied up front. Benchproofer converts real repository issues with a known correct patch into formally verified tasks, yielding SWE-Proof: all 500 SWE-bench Verified instances with ground-truth specifications, implementations, and proofs under Nagini, Velvet, and Lean.
- The pipeline formalizes only the new or modified code, summarizes unchanged callees with fuzz-tested axioms, and admits an instance only after mechanical gates (verification, rejection of the pre-fix code, mutation kill rate, hygiene) and adversarial LLM audits (axiom soundness, conformance on at least 10^5 inputs, soundness/completeness, equivalence, leakage) all agree; it also built gate-passing artifacts for 242 of 266 Python tasks in
SWE-bench Pro. - Hidden tests prove substantially incomplete: an adversarial audit drops the unaided baseline from 85.0% to 58.2% for
Claude Opus 4.8and from 81.2% to 33.4% forGPT-5.5, and a structured natural-languageEARSspecification does not close the gap, whereas formal verification against the ground-truth specification costs under one point. - Supplying the ground-truth formal specification lifts resolution to roughly 95% for both models, about +7 to +8 points beyond localization alone, while agents asked to write and verify their own specification gain nothing over the baseline (best cell +0.6, worst -1.2).
- Specification synthesis is the bottleneck: only 46–72% of model-written specifications pass a five-property audit, with faithfulness (modeling too little of the behavior the issue requires) failing on 42.5% for Opus and 30.0% for GPT, and specifications from unresolved instances failing the audit 89.4% of the time versus 47.3% for resolved ones.
- The verify-implies-resolve guarantee is not itself a proof, since it rests on audited trust points (axiom soundness and the agent-written Python translations used for conformance under
VelvetandLean), the auditors and judges are LLMs whose verdicts are evidence rather than proof, and patches that only rename, relocate, or change side effects fall outside what pre- and post-condition specifications can observe.
Calibrating Teacher--Student Discrepancy for On-Policy Distillation
On-policy distillation (OPD) trains a reasoning model on the token-level discrepancy between a stronger teacher and the student's own samples, but that discrepancy also contains deviations arising from the teacher itself, a problem amplified when the teacher is given privileged information. Calibrated On-Policy Distillation (Cal-OPD) estimates the teacher's self-deviation region using positive and negative privileged interventions and keeps only the part of the discrepancy lying beyond that region. On mathematical reasoning benchmarks, it consistently outperforms standard OPD and its variants across model scales while retaining only about 52-65% of the original discrepancy as the optimization signal.
On-policy distillation (OPD) trains a student on the token-level log-likelihood gap to a teacher, but that gap also carries "teacher self-deviation" (TSD) — shifts in the teacher's own likelihoods under changed context that reflect no task knowledge — and privileged-information OPD amplifies it. Cal-OPD probes the teacher with a positive and a negative context intervention to estimate a per-token TSD region, then distills only the part of the discrepancy that falls outside it.
- Rescoring ~60M tokens of fixed
Qwen3-1.7Brollouts with aQwen3-8Bteacher shows TSD needs no task knowledge and mostly ignores correctness: task-agnostic instructions shift 29.8% of tokens versus 29.4% for answer-level hints, and swapping a correct answer for a wrong one still gives 88.7% directional agreement in the shifts. - TSD concentrates on surface-form tokens rather than mathematical content, with discourse markers such as
maybe,however, andthereforeshifting in over 89% of occurrences while digits, symbols, and LaTeX notation stay below 9.3%. - The method takes the largest upward and downward teacher log-likelihood shifts under the two interventions (evaluative feedback praising or condemning the rollout), widens that interval by a relaxation factor
λ=5, and sets the advantage to zero when the student's log-likelihood lies inside it, otherwise keeping only the residual beyond the nearest boundary. - On six math benchmarks (Avg@16)
Cal-OPDaverages 53.1 forQwen3-4B-Thinking-2507→Qwen3-1.7Band 69.0 forQwen3-30B-A3B-Thinking-2507→Qwen3-4B, which is +2.3 and +3.1 over standard OPD, while standard OPD drops the 4B student below its starting 66.6 andPrivileged-OPDis the worst distillation method in both settings. - It retains only 52–65% of the original discrepancy, zeroes out 27–34% of tokens, and ends training at about 9.3K-token responses versus 11.8K for OPD, making it roughly 1.26× faster despite the extra teacher passes.
- Results are sensitive to the probe and to
λ: calibrating with solution-level interventions over-filters to about 20% retained signal and falls to 49.0, largeλends up below theλ=1setting, and evidence is limited to the Qwen3 family, math benchmarks, and 100 training steps.
GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills
Skills give large language model (LLM) agents task-specific procedural guidance, but writing them as unstructured natural language leaves them redundant, short on workflow-level direction, and hard to optimize over a huge search space. The authors represent a skill as a graph whose nodes are execution steps with operational guidance and whose directed edges encode context-dependent transitions. GraphSkillEvo then runs population-based evolutionary optimization over these graphs with mutation and crossover operators. Across five agent benchmarks it beats the SkillOpt baseline, improving average accuracy by 4.01% on GPT-5.4-nano and 1.76% on GPT-5.4.
Skills that guide LLM agents are usually optimized as unstructured natural-language text, which agents find hard to follow and which gives the optimizer a large, redundant search space. GraphSkillEvo instead represents a skill as a graph, where nodes are execution steps with their own guidance and edges are condition-dependent transitions, and evolves a population of such skills with structure-aware mutation and crossover.
- A skill consists of global guidance, reusable step nodes, and workflows that pair an applicability condition with an ordered node path. A population of 4 skills evolves for 5 generations using four round-robin operators (mutation and crossover on either the guidance or the graph), where mutation reflects on up to 5 failed trajectories from 15 sampled training instances, a validator script rejects malformed graphs, and the top 4 by validation score survive.
- Across five benchmarks (
SearchQA,SpreadsheetBench,DocVQA,LiveMathematicianBench,ALFWorld) it is best in 13 of 14 settings, beatingSkillOpton average by +4.01 points onGPT-5.4-nano, +1.76 onGPT-5.4, and +1.33 under the Codex harness, with the largest gain onSpreadsheetBenchwith the nano model (60.71 vs 50.11). - Total optimization cost is lower, since
SkillOptuses 1.31x the tokens ofGraphSkillEvoonGPT-5.4and 1.36x on nano, althoughGraphSkillEvospends more onALFWorld(47.75M vs 28.18M tokens on nano). - In ablations on nano, the three-benchmark average falls from 71.52 to 64.08 without the graph structure, 66.59 without crossover, and 54.50 without mutation, while skills optimized on nano and transferred to
GPT-5.4reach 71.78 onSpreadsheetBenchversus 69.40 for direct optimization and 53.21 for transferredSkillOpt. - The gains are modest on the stronger model and it trails
SkillOptby 0.80 points onLiveMathematicianBenchwith nano. In the five-run significance test onlySpreadsheetBench(p=0.016) andALFWorld(p=0.038) clear 0.05, run-to-run variance is often higher thanSkillOpt's (SearchQAstd 3.59 vs 1.25), and everything is evaluated on a single model family with validation sets as small as 18 instances.
A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
A language model may hold knowledge it does not report, for example by sandbagging on a capability evaluation, and its outputs alone cannot show whether it is hiding an answer or lacks one. Borrowing the forensic Concealed Information Test, Probe of Internal Recognition (PIR) presents a question with candidate answers and reads from the model's internal states which candidate it recognizes as correct, without needing an honest reference model or a labeled truth corpus. Across eight models from five families (Gemma, Qwen, Llama, Mistral, Phi), PIR recovers the recognized answer at 0.70 to 0.87 balanced accuracy, against a 0.28 to 0.40 unknown-item baseline and 0.25 chance, and recognition stays between 0.85 and 0.93 under prompted deception, trained sandbagging, password-locked checkpoints, and circuit-broken checkpoints. When unlearning actually removes the knowledge, recognition falls to the level of a never-known question, so the probe separates a model that will not answer from one that cannot; the signal is reported to be causal, to add information beyond black-box cues, and to extend to free-form generation.
A language model's wrong answer is ambiguous — it may be sandbagging or may genuinely not know — and outputs alone cannot tell the two apart. PIR (Probe of Internal Recognition) adapts the forensic Concealed Information Test to activations: it presents a question with its candidate answers and reads which one the model internally recognizes as correct, with no honest reference model or labeled truth corpus.
- The method averages, over questions the model answers correctly, the difference between the correct option's last-token hidden state and the mean distractor state, then scores each option of a new question by projection onto that direction, with a label-free divergence detector comparing the recognized option against the one the model actually output.
- Under prompted deception, recognition of the concealed answer reaches 0.70–0.87 balanced accuracy across eight models from five families (
Gemma,Qwen,Llama,Mistral,Phi) against an unknown-item baseline of 0.28–0.40, and it holds at 0.85 on a trained sandbagging policy and 0.93 / 0.90 on external password-locked and circuit-brokengemma-2-9bcheckpoints. - Unlearning instead erases the signal:
RMUongemma-2-9bdrops recognition from 0.85 to 0.39, overlapping the 0.35 baseline, and externalQwen2.5-7Bcheckpoints fall from 0.85 to 0.36 (RMU) and 0.38 (NPO), so the same readout separates a model that will not answer from one that cannot. - The direction is causal and adds to behavioral cues: steering along it makes
gemma-2-9bpick the steered wrong option 18.9% of the time versus 0.5% for random directions and 0% for an output-direction control, while combining it with black-box features lifts detector AUC to 0.93 and 0.90 on the external organisms, and a free-form variant onTriviaQAreaches AUC 0.776. - An organism adversarially trained to zero the recognition margin drives the readout to baseline while keeping 0.71 capability (the authors retract an earlier in-sample claim to the contrary), and the method also needs enumerable candidate answers, requires the model's base checkpoint for deployment detection, cannot tell erased knowledge from knowledge never held, and rests partly on single-seed organisms and one clean
RMUrun.
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Computer-use agents have developed along two separate lines, graphical interface control and coding through the command line, while real digital work interleaves both. RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web where an agent is given a running reference application, must discover its behavior, and must build a faithful reimplementation, with the reference serving as an oracle for hidden behavioral tests that yield execution-grounded rewards. Models trained on the generated trajectories improve on five out-of-distribution coding and hybrid computer-use benchmarks and verify their rendered outputs more often. On the held-out 250-task RecreationBench, GPT-6 Astra leads at 58.1% overall but passes all programmatic tests on only 2.8% of tasks, with agents reproducing static interface structure more reliably than interactions and computed outputs.
GUI agents can operate software but cannot build it, while terminal agents can build software but cannot see the interface they produce. RecreationWorld trains and evaluates hybrid computer-use agents on recreation: given only a running reference application, the agent must explore it, implement a faithful clone, and visually verify its own build. The reference serves as the oracle for hidden behavioral tests.
- The framework provides isolated, reproducible environments on Ubuntu, macOS, Windows, Android, and Web, with native GUI control plus coding tools and a 20-hour rollout budget. Candidates are scored by frozen programmatic assertions (via
AT-SPI,AXUIElement,UI Automation,UiAutomator, or DOM) and by visual assertions judged byQwen3.7-Plus, all validated on the reference and by human reviewers. - For training,
Qwen3.8-Maxrollouts on open-source apps were rejection-sampled into a balanced 35,000-trajectory SFT mixture (7,000 per platform). Fine-tuningQwen3.7-PlusandQwen-Flash-CPTon this mixture left both above their first checkpoint on all five out-of-distribution benchmarks (ProgramBench,GameCraft-Bench,Vision2Web,OSWorld 2.0,WeaveBench), with gains of up to 17.9 points and more frequent checking of their own rendered output. - On the held-out
RecreationBench(250 tasks, 50 per platform, median 282.5 tool calls per trajectory),GPT-6 Astraleads ten frontier models at 58.1% overall, ahead ofClaude Opus 5at 44.2% andGPT-5.6 Solat 42.1%. It passes every programmatic test on only 2.8% of tasks, and no other model exceeds 0.8%. - Agents reproduce static interface structure more reliably than interactions and computed outputs. 89.4% of recreations are smaller than their reference (median 16.9% of reference LOC) and more monolithic, and fewer than half of trajectories relaunch and inspect the app after their final edit (23.6–47.5% across models).
- The trajectory and behavior analyses are descriptive rather than causal, and the transfer curves use one trial per task and are not monotonic. Web tasks cannot be source-blind and Android lacks a packet-level egress filter, while agents repeatedly attempted network and protected-path access, so the benchmark's integrity rests on enforced isolation rather than prompt compliance.
Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention
Gating the value pathway of attention is reported to improve language model pretraining, but prior studies disagree about why. The authors argue that such gates supply two things softmax attention lacks: abstention, which lets a head output nothing despite attention weights having to sum to one, and noise filtering, which suppresses interference from superposed features in the residual stream. In matched models from 10M to 350M parameters, abstention is supplied through a learned per-head sink logit and filtering through a gate on each value, and the benefit of abstention shrinks with scale while the benefit of noise filtering grows, with abstention accounting for nearly all of the gain at 10M and filtering for most of it at 350M. The best model at every scale has both primitives, which add negligible parameters and remain compatible with the key-value cache.
Softmax attention forces every head to spend its full attention mass somewhere and aggregates value reads linearly, so a head can neither output nothing nor suppress interference from superposed features. The author argues that value gating helps because it partially supplies two separate primitives, abstention and noise filtering. Each is isolated with its own mechanism, a learned per-head sink logit for abstention and a per-value gate for filtering, in matched models from 10M to 350M parameters.
- Abstention is a phantom key with a learnable logit and a value fixed at zero, while filtering is either a
norm gate(a threshold on read norm) or aprojection gate(a sigmoid of a learned linear projection of the value). A lemma shows that a value gate equals a routing gate plus a partial zero option, which motivates a routing-slotrenormcontrol. - On
FineWeb-Eduwith paired seeds, the sink logit's gain over baseline shrinks from 0.0185 nats at 10M to 0.0013 at 350M, while the filtering gain ofcombo2over the sink logit grows from 0.0036 to 0.0114, so the two curves cross between 50M and 124M. - The abstention benefit shrinks because larger baselines imitate it with heavier attention sinks (11% of mass on position 0 at 350M versus 6.0% at 124M), even as the sink-logit model's phantom absorbs 13% to 53% of attention mass.
- The two benefits are largely additive, and
combo2(sink logit plus projection gate) has the lowest loss at every tier (3.0210 vs 3.0338 at 350M), costs under 0.01% extra parameters, and stays compatible with the key-value cache. - Injecting transplanted value reads shows that each filter is blind along its own decision variable: the
norm gateloses 4.4 nats vs the baseline's 2.6 at dose 0.8, and theprojection gatecollapses only under junk aligned with its learned direction, whereascombo2loses 1.1 vs 5.9 nats at the highest dose at 350M. - The 350M results are single-seed, all training uses one corpus at modest token budgets, the effects are a few thousandths of a nat with no transfer tested beyond 350M, and no throughput was measured.
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
Reinforcement learning for coding agents needs diverse tasks with reliable verifiers, and existing pipelines mine them from development artifacts such as issues and commits, which limits what can be extracted. CodeMidas is an agentic pipeline that uses source code as its only task-specific input: agents explore implemented functionality to write behavioral specifications, construct tests grounded in executing the original code, and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on it with GRPO improves all five benchmarks tested, including DeepSWE +11.7%, ProgramBench +17%, and Terminal-Bench v2.1 +8.5%, and the trained agent explores codebases more and self-verifies in more varied ways.
RL for coding agents needs many diverse tasks with trustworthy verifiers, but existing environment pipelines depend on issues, pull requests, commits, existing tests, or documentation, which limits what can be extracted from a repository. CodeMidas is an agentic pipeline that builds an executable RL environment from source code alone: it removes an already-implemented feature, has the solver rebuild it from a behavioral specification, and grades the result against hidden tests grounded in running the original code.
- An agent selects functionality with public entry points (CLI tools, pure functions, stateful APIs), deletes the core implementation to form a coherent starting codebase while keeping the original as the reference solution, then builds tests from reference execution and reviews every assertion to remove restrictions the statement does not justify, such as exact message wording or incidental ordering.
- Candidates must show a clean fail-to-pass transition across six fresh containers (two starting-state runs fail, four reference runs pass), then survive adversarial leakage-hunting rollouts, an agent review of four solution attempts for verifier false positives and negatives, and a filter that keeps only tasks with mixed outcomes under a frontier model, which leaves 5,545 tasks from 3,185 codebases across 23 languages and 15 domains, with a median reference patch of 142 lines.
- Training
MiMo-V2.5withGRPOand binary execution rewards improves all five external benchmarks, includingDeepSWEfrom 10.0% to 21.7%,ProgramBenchAlmost Solved from 4.5 to 21.5, andTerminal-Bench v2.1from 63.7% to 72.2%, while the held-outCodeMidas Valpass rate rises from 35.0% to 44.7%. - Scaling the filtered pool from 1k to 3k to 5,545 tasks lifts
DeepSWEfrom 17.57 to 19.05 to 21.70, and even the 3k filtered subset beats an unfiltered 8k sample onSWE-bench Pro,DeepSWE, andCodeMidas Val; behavior also shifts, with pre-edit read/search calls rising from 27.2 to 40.1 and agent-written checks associated with a 4.2-point higher pass rate (95% CI 1.8–6.6). - The evidence comes from a single base model with no comparison against training on other environment datasets, the cleaning and filtering stages are ablated only jointly, the full pool's edge over the unfiltered 8k sample on
SWE-bench Prois just 0.59 points, and the behavioral links are correlational, with confidence intervals for exploration and drafting spanning zero.
Applications 62
From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators
Existing evaluations of large language models in medicine target static outputs rather than whether a patient actually understands a discharge plan after an interactive conversation. DischargeBench simulates multi-turn sessions in which a candidate LLM educator teaches a persona-driven Virtual Patient, while an Education Monitor Agent keeps the patient realistic without altering the educator. The accompanying MIMIC-IV-Ext-DischargeBench dataset has 477 cases across 24 ICD chapters with persona axes for personality, education level, health literacy and medical-history recall, scored on four axes by an LLM judge aligned to physician annotations. Across closed- and open-source models, aggregate scores hide clinically relevant variation across conditions and personas, with difficult personas exposing coverage failures, comprehension gaps and lower factual consistency.
Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR
Automatic speech recognition (ASR) systems tuned for Word Error Rate (WER) often miss named entities and filled pauses in accented conversational English, both of which matter for language-learning feedback. The pipeline targets speakers from India, Indonesia and Latin America in three stages: heuristic SQL filters that curate training data with 2.8x the entity density of random sampling, regional LoRA adapters on Qwen2.5-Omni-3B that emit verbatim and corrected transcripts in one forward pass, and a six-category error taxonomy checked by an LLM judge. It reaches 80-85% entity recall (up from 53-55%) and 76-86% filler recall (up from under 5%) at 6-10% WER, beating Whisper and a commercial ASR on entity recall and matching a zero-shot 30B model with ten times fewer parameters. Paired bootstrap tests attribute 2.8-4.2 percentage points of the entity recall gain to data curation alone.
Towards Secure Cloud-Native Computing: Unveiling Kubernetes Misconfigurations with Large Language Models
Kubernetes misconfigurations are a common source of security and performance problems in cloud-native deployments, and this study examines how well Large Language Models (LLMs) can detect them. The authors build a taxonomy of common Kubernetes misconfiguration types, empirically benchmark existing state-of-the-art detection tools, and analyze which Kubernetes objects are most prone to misconfiguration and how severe the resulting issues are. The abstract reports no specific detection accuracy figures, presenting the work as insight into how LLM-based detection can complement existing methods.
MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery
Symbolic regression recovers closed-form equations from data, but search-based methods rely on costly combinatorial optimization from random starts, while pretrained neural models produce formulas quickly yet often with symbolic errors. MOSAIC-SR uses a pretrained Transformer to propose multiple initial expression sketches that seed searches in several promising regions, with each search jointly recovering structure and constants through scale-aware constant optimization and local symbolic repair. Evaluated on SRSD-Feynman with and without dummy variables plus six additional benchmarks, it obtains the highest symbolic solution rate on every dataset while ranking in the top two for predictive accuracy, and the advantage holds when irrelevant dummy inputs are present.
How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?
Language models can write GPU kernels that beat PyTorch, but how much of a real workload those kernels affect has not been measured. On KernelBench level 1, a frontier model produces correct kernels for 91.1% of problems with verified speedups on 22 of 56, while the best open-weights model reaches only 30.4% correct. Profiling seven workloads shows the addressable fraction of wall-clock time ranges from 8.9% to 58.2%: on transformers, 80-86% of runtime sits in cuBLAS matrix multiplies and FlashAttention, bounding realistic end-to-end improvement at roughly 1%, whereas recommenders are far more addressable, and the new DLRM-Bench projects an 8.63% end-to-end gain there. The authors also find that KernelBench's torch.allclose correctness check accepts an all-zeros tensor on 4 of 60 problems, which two of their own kernels exploited, including one scored at 283x, and they propose scale-invariant replacements.
Talk to Me, Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars
Cloud-hosted language models bring network dependency and variable latency, which makes them a poor fit for voice control in time-critical autonomous driving. Jarvis is an offline, open-source voice assistant for issuing high-level behavioral commands to autonomous racecars, combining speech recognition, speech synthesis, and a text-to-command classifier built by domain-specific fine-tuning of Mistral 7B. It reaches 97.63% intent recognition accuracy with an average processing latency of 1.39 s, outperforming larger online-hosted models in the authors' evaluation.
From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost
Common measures of AI productivity capture what output was produced but not the interaction effort needed to get there. Drawing on economics, the authors propose evaluating human-AI collaboration as outcome quality relative to interaction cost and apply it to two datasets spanning four tasks. They find that sessions with identical quality ratings can differ by up to 70 times in interaction cost, that quality-cost relationships vary by task, and that subjective user ratings are not reliable proxies for productivity. Productive sessions are marked by agents probing earlier and users spending less effort repairing the interaction.
Scaling Forced Alignment to End-User Devices
Forced alignment of audio to text with the Viterbi algorithm costs quadratic time and memory in many implementations, which rules out long recordings on ordinary hardware. Two optimizations are proposed: the Hirschberg algorithm to align in place using linear memory, and modeling speech-to-text alignment as a constrained random walk so the search space can be pruned with a tunable confidence bound that still tolerates transcription errors. Memory for a three-hour input drops from 140 GB to 5 MB while producing identical alignments in a third of torchaudio's CPU time, and pruning adds a further 2x speedup on inputs longer than 20 minutes with accuracy preserved in over 98% of tested cases.
Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake
Health systems weighing AI-assisted psychiatric intake need a repeatable way to check these tools against clinical standards without consuming much clinician time or assuming one interviewing style. InterviewPlayground supplies a memory-augmented simulated patient built from expert-authored vignettes, together with a simulated intake platform and evaluation modalities tailored to the task. In a pilot of six clinicians running 25-minute assessments against a GPT-based interviewer, the language model recovered 88.0% of the clinically relevant vignette items versus 38.9% for clinicians, but made unsupported clinical inferences more often (56.8% versus 27.8%) and characterized identified safety concerns less often (33.3% versus 66.7%).
SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity
Off-target protein binding causes many small-molecule side effects, yet structure-based design mostly generates selective compounds from scratch instead of improving drugs already well characterized. The proposed specificity optimization task asks for constrained edits to an existing compound that widen its preference for the intended target over measured off-targets, scored on a new ChEMBL-derived benchmark; an agent docks each compound against all targets, compares poses through residue-aware atom-protein contacts, and hands those differential interactions to a language model that proposes modifications, keeping only candidates passing similarity, ADMET, and docking-selectivity filters. Across 915 compounds the agent widened the target/off-target binding gap for 84.8%, shifting the mean from -0.72 to +0.47 kcal/mol at mean Tanimoto similarity 0.72, and ablations show that replacing residue identities with binary contact flags eliminates the improvement entirely.
Identifying Security Platform Product Abuse with Machine Learning
Sophisticated threat actors can misuse security platforms inside customer environments or run bypass experiments against the product itself, often using living-off-the-land (LOTL) techniques rather than easily detected malware. The authors describe a deployed machine learning detection system for this kind of product abuse, designed around multiple data modalities spread across different databases, a cold-start problem caused by the rarity of such events, and operational limits on cost and performance. They report a 35% increase in product abuse coverage alongside a 30% reduction in monthly alerts, plus adaptability to changing attacker behavior. A retrospective evaluation examines the value of explainable features and counterfactual performance on previously identified attacks.
Fast And Accurate Text Content File Type Identification
Identifying file types from content matters in cybersecurity, where magic numbers and extensions cannot be trusted, but model-based tools such as Magika are computationally heavy and parser-based tools are less accurate. The authors propose a neural network specialized for text-content files, especially source code. On open-source files it is more accurate on average than existing tools while being approximately four times faster than Magika and 28% smaller.
Co-Evolving Zero-Day Jamming: Adaptive Attack Synthesis and Graph Attention-Based Online Detection
Detectors for previously unseen (zero-day) radio jamming strategies are hard to evaluate because existing attack models assume prior knowledge of the target receiver, and existing detectors miss the global temporal-spectral structure of jamming and cannot separate new strategies as they emerge. The authors pair an online detector, which combines a graph attention network (GAT) with Dirichlet process (DP)-means clustering to classify known strategies and discover new ones, with a reinforcement learning jammer that treats the receiver as a black box, infers the detector's state through hypothesis testing, and trades attack impact against stealth. In simulation the jammer achieves 33% higher attack efficacy and 67% higher stealth than benchmark attackers, and the detector reaches 20% higher detection accuracy than benchmark detectors against it.
Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis
Farmer.Chat, an advisory service for smallholder farmers, must diagnose crop problems from a single low-quality phone photograph with no accompanying text, and its production system offers no adjustable thresholds or extensible label set. An analysis of about 1.16 million photographs from Ethiopia, India, Kenya, and Nigeria shows the production quality gate rejected 46.8% of images and that 35.8% of problems labelled as disease were actually pests. The work splits diagnosis into a quality gate, a crop detector, and a disease or pest detector, comparing a single fine-tuned Qwen3-VL-4B against small specialists (DaViT, YOLO26, MobileNetV3). A MobileNetV3 gate replaces GPT-4o at 86.9% F1 in 12 ms, and a hierarchical DaViT-Base reaches 95.41% crop accuracy versus 91.46% for the production baseline, while the fine-tuned vision-language model uniquely handles all stages in one call and can ask for a better photo.
Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening
Benchmark scores for medical vision-language models may not hold up when the evaluation setup changes, so the authors audit BioMedCLIP, CheXficient, MedSigLIP, and a general-domain OpenCLIP comparator for tuberculosis screening on 12,200 chest radiographs from Montgomery, Shenzhen, TBX11K, and VinDr-CXR. No model leads on every cohort and reliability criterion, prompt wording changes AUROC in 21 of 48 controlled comparisons, and replacing healthy controls with sick non-tuberculosis controls lowers AUROC by 0.075 to 0.306 for all four models. Thresholds set for 95% sensitivity on TBX11K hold in only four of sixteen target evaluations, and a supervised model drops from 0.999 AUROC on validation to 0.629 on external cohorts. The authors conclude that portability claims should name the full evaluation specification instead of being attributed to a checkpoint.
LLM-Generated Feature Pools for Time Series Anomaly Detection
The authors ask how far a simple statistical pipeline can go on univariate time series anomaly detection: compute a small pool of sliding-window statistics, score windows with a robust median-absolute-deviation model, and pick a feature subset per domain on a held-out split. On TSB-AD-U it reaches 0.529 per-series VUS-PR, above the best neural (0.45) and statistical (0.44) leaderboard entries, and ablations show the candidate feature pool moves the score far more (0.226) than the selection strategy (0.031). They therefore generate a pool per domain by prompting a multimodal LLM with example windows from that domain. Generated pools match the hand-crafted one, and selecting over the union of both lifts the pipeline to 0.588, matching the best leaderboard entry.
EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise
Enterprise generative AI projects often fail to show business impact, which the authors attribute largely to measurement: public benchmarks show what a model can do, not whether a specific workflow is reliable, safe, and worth scaling on an organization's own data and controls. EnterpriseVal is a use-case-level evaluation system comprising a formal specification of the frozen configuration under test, a metric catalogue, a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference, an executable threshold gate producing REJECT/CONDITIONAL/SCALE decisions, and a value-and-risk model. In a pilot at a global bank, credit-memo drafting reached 88% citation precision and a 1.6% hallucination rate against gates of 70% and 5%, and analyst refinement effort in procedure transformation fell from an estimated 27.4 to 2.9 hours per document. The authors explicitly separate established results from pilot evidence and open hypotheses.
LLMs as Feature Engineers for Text-and-Tabular Prediction
The framework automates extraction of interpretable, schema-bound categorical features from unstructured text for use in tabular prediction models. A generator LLM proposes semantic feature definitions, a separate extractor LLM materializes them, and a downstream tabular model scores them, with explicit model errors such as AUC ranking inversions translated into natural-language feedback that steers the next round of proposals. On three public datasets this error-driven loop speeds up feature discovery by up to 3x compared with unguided search, and the generated features complement TF-IDF and dense embeddings so that the combination beats any subset. The discovered features also dominate SHAP importance rankings, giving a semantic audit trail for each prediction.
44 more specialized papers
- ResNLS: An Improved Model for Stock Price Forecasting Yuanzhe Jia, Ali Anaissi, Basem Suleiman
- A Hybrid Computational Intelligence Framework for scRNA-seq Imputation: Integrating scRecover and Random Forests Ali Anaissi, Deshao Liu, Yuanzhe Jia et al.
- HERMES: Contrast-Aware Knowledge Graph Reasoning from Clinical Notes for Patient Outcome Prediction Gia-Bach Nguyen, Hoang-Ha Nguyen, Tuan-Cuong Vuong et al.
- SAGE: Schema-Guided LLMs for Grant Review Erik Varapaev, Andrei Chetvergov, Stepan Ukolov et al.
- Transsion's Speaker-Attributed Multilingual ASR System for the MLC-SLM 2026 Challenge Zhecheng Ren, Xuanji He, Xiaoxiao Li et al.
- Generative inversion for early ranking of competing geologic interpretations Harun Ur Rashid, Daniel O'Malley
- Trustworthy FinAInce: Unpacking How AI-Mediated Financial Advice is Judged Aryan Ramchandra Kapadia, Eshwar Chandrasekharan, Koustuv Saha
- From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities Desta Haileselassie Hagos, Saurav Keshari Aryal, Legand L. Burge
- Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation Mohit Chandra, Nabin Kim, Eli Min et al.
- Signal-Centric Remote Sensing via Alternative Preprocessing and Acoustic Processing for ML-Driven Applications Logan Luna, Sirio Jansen-S\'anchez, Ilteris Demirkiran et al.
- EnSol: an environment-aware graph neural network for molecular solubility prediction Thao Nguyen, Saman Shafaei, Zhengyi Zhang et al.
- Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization Lingfei Kong
- Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis Boyuan Zhao, Meng Ye
- MIRCID: Inferred Hub-miRNAs Drive Cross-Task Improvements in Drug Mechanistic Modeling Xin Cao, Yigang Chen, Jiatong Xu et al.
- Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale Hao Fu, Jichao Sun, Baiting Zhu et al.
- Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding Chenqian Le, Beatrice Fumagalli, Yasamin Esmaeili et al.
- Routine Blood Tests Outperform CRP for Distinguishing Bacterial From Viral Infection in Children Mihaela Demireva, Zhecho Mitev, Djuna Chinareva-Klimentova et al.
- Knowledge-Graph-Augmented Chronos-2 for HEC-RAS Surrogate Forecasting Edward Holmberg, Elias Ioup, Mahdi Abdelguerfi
- Probabilistic Forecasting of Business Process Executions with Neural Temporal Point Processes Jiaxin Yuan, Daniela Grigori, Han van der Aa
- Consistent Relexicalization of Clinical Documents using Graph-Based Approach Dipankar Das, Atri Mandal, Sandeep Singh et al.
- Offline Multimodal Large Language Models for Decision Support in Air Operations Joao P. A. Dantas, Jelton A. Cunha, Gabriel Dietzsch
- Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation Kensei Nosaka, Shunnosuke Ikeda, Yuichi Takano
- Interference-Driven Clustered Optimisation for FM Spectrum Coordination Federica Mangiatordi, Emiliano Pallotti
- Efficient Architecture Search under Leave-One-Subject-Out Evaluation Heinke Hihn, Friedhelm Schwenker
- Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks Giambattista Amati, Federica Mangiatordi, Pierpaolo Salvo et al.
- Dual-Interest Sequential Product Recommendation With Multi-Granular SSM Shuiying Liao, P. Y. Mok
- Periodic Neural Mapping for Unsteady Rotor-Blade Pressure and Aeroelastic Load Prediction Lionel Salesses, Joachim Dominique, Tariq Benamara et al.
- Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection Syed Ali Ahmed (National University of Computer and Emerging Sciences, Karachi, Pakistan) et al.
- Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education Andy Gray, Jake Hobbs
- Bayesian classification of astronomical spectra with class uncertainties Simon Barton, Martin Sahl\'en, Andreas Korn et al.
- Bilevel Optimization of Topology and Hyperparameters (BOTH) Suryanarayanan Manoj Sanu, Miguel Anibal Bessa, Alejandro Marcos Arag\'on
- Complete Neural Electronic Initialization Accelerates Materials DFT Felix {\AE}rtebjerg, Jonas Elsborg, Arghya Bhowmik
- Per-Aetiology Contrastive Severity Embeddings with Phonological Pseudo-Labelling for Multilingual Dysarthric Speech Bernard Muller, Antonio Armando Ortiz Barra\~n\'on, LaVonne Roberts
- RegKT: Interpretable and Robust Deep Knowledge Tracing With IRT-Regularizer Samuel Girard, Juan D. Pinto, Jill-J\^enn Vie et al.
- MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention Muhammet Sami Yavuz, Sabri Mustafa Kahya, Richard R. Chen et al.
- Adaptive Uncertainty-Aware Modeling and Stochastic Radial Basis Function Predictive Control for Personalized Fluid Resuscitation Elham Estiri, Hossein Mirinejad
- Touvigation: Embodied Adaptive Object Acquisition for Blind and Low-Vision Users in Unfamiliar Indoor Environments George Xi Wang, Xiangyu Li, Shaoyue Wen et al.
- Chronosphere: Space-Time Tessellation of Local Climate Experts Daniel Cher, Eric Xing, Kexing Li et al.
- Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models Fangzhou Wang, Yixuan Yang, Camilla Balzarotti et al.
- Joint Remaining Useful Life Prediction and Capacity Estimation of Lithium-Ion Batteries Using Partial-Charging Data Khoa Tran, Ho-Si-Hung Nguyen, Phone Wai Yan Moe et al.
- Learning Cardiac Features: ECG Biometrics Across Time and~Exercise Luca Thiebaud (AMU, AMU SCI, DIAPRO et al.
- Assessment of Machine Learning-Based Critical Heat Flux Models in the CTF Subchannel Code for Square Rod Bundle Prediction Aidan Furlong, Vinicius de Melo Monteiro, Robert Salko et al.
- BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings Alexandre Andre, Shivashriganesh P. Mahato, Vinam Arora et al.
- Cross-sector generalization of accident-process role classification in occupational accident narratives Aho Yapi, Pierre Latouche, Arnaud Guillin et al.
Large Language Models 39
Do small language models know what they don't know?
Entropy-based confidence signals are tested as a way to improve Small Language Models (SLMs) under 3 billion parameters running entirely on consumer hardware, using seven approaches across 7 model pairs and 5 natural language understanding benchmarks. Token-level entropy turns out to be effectively blind at this scale: in 91% of dataset-model combinations, mean token entropy is near zero whether or not the answer is correct. Semantic entropy, computed by sampling several answers, clustering them by meaning and measuring the spread, does recover a usable signal, and routing uncertain queries to a larger expert model raises accuracy by up to 50 percentage points. Cross-family routing such as SmolLM 360M to Phi-3.5-mini averages +22.0% versus +6.8% for same-family routing, indicating that expert quality matters more than architectural compatibility.
Reviser: Revision-Capable Text Generation via Autoregressive Cursor Actions
Text generators that can go back and revise earlier content usually pay for that flexibility with repeated computation over the whole sequence. Reviser is a decoder-only Transformer that writes onto a mutable canvas by predicting one cursor-relative action per step, either INSERT(token), MOVE(Δ) or STOP, so it is autoregressive over the edit history rather than over final text order. On a continuation benchmark it is strongly preferred to the diffusion-style models SEDD and MDLM in arena evaluations, and trajectory statistics show it genuinely makes frequent backward moves and mid-canvas insertions. It is competitive with size-matched autoregressive baselines at 100M and 300M parameters and needs substantially less inference compute than multi-pass refinement and diffusion-style baselines.
When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
As model-written peer reviews enter training corpora, later AI reviewers may learn from earlier ones, and this study simulates one step of that loop by fine-tuning Llama 3.1 8B on official ICLR reviews from 2018-2023 and then training four successors on ICLR 2024 data with varying mixes of official and model-generated reviews. Adding synthetic reviews compresses rating distributions and reduces semantic diversity both within a paper's reviews and across the corpus, a pattern the authors name scientific-judgment collapse. To counter it they introduce TrustReviewer, an open-source reviewing system that trains in a single stage on a corpus curated to remove low-quality and degenerate supervision, and applies paired activation steering at test time to correct residual collapse without further training or expert annotation.
RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
Long-context inference is increasingly bottlenecked by prefill, where dense self-attention processes the whole prompt before generation starts, and block-sparse methods that score blocks by their centroid can miss a single highly relevant token buried among irrelevant ones, a failure the authors call mean dilution. RBS-Attention is a training-free sparse-prefill method that combines a centroid branch for average relevance with a rescue branch that uses each key block's maximum radius and its prompt-, layer-, and head-dependent distribution to flag blocks at risk of underestimation, while keeping regular block-sparse FlashAttention execution. On H100 GPUs at 128K context with Qwen3-30B-A3B-Instruct-2507-FP8 it achieves 20.65x standalone prefill-attention speedup, 11.92x within vLLM, and 5.97x end-to-end time-to-first-token speedup. On dense Qwen3-32B it scores 88.65 overall on RULER versus 89.52 for dense attention, with further evaluation on LongBench-v2, InfiniteBench, and Video-MME.
Attention-Aware Routing: Coupling Routing and Attention in MoEs
Routers in Mixture-of-Experts language models usually choose experts from a token's hidden state alone, using little contextual information. Attention-Aware Routing (AAR) augments the router with temporal and spectral features extracted from a sliding window of attention weights, and trains only the routing parameters while the base transformer stays frozen. On OLMoE it improves GSM8K by +3.37 percentage points over a routing-only supervised fine-tuning baseline, and shortens incorrect, long-diverging generations while leaving correct answers' length unchanged. The authors also show routing and attention form a coupled circuit, where routing changes at one layer amplify attention sinks at the next, and that the method is depth-sensitive: applying it at all layers can hurt factual retrieval, whereas math reasoning gains persist when it is applied deeper in the network.
Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing
Hallucinated and faithful responses are distinguished here by analyzing the topology of information flow in a model's attention graphs. The method uses Forman-Ricci curvature to find structural bottlenecks and combines semi-local and global flow features of attention heads, requiring only a single forward pass. It consistently improves over existing attention-based and multi-response baselines on two hallucination-detection benchmarks across several model architectures. Further analysis ties hallucination to impaired context sharing among tokens: over-reliance on self-attention, diffuse retrieval from earlier tokens, or information over-squashing, especially in the final transformer layer.
Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models
How fine-tuning reshapes a language model's internals is examined by comparing changes in attention patterns and layer-wise activations against the task-relevant components identified by EAP, a circuit-attribution method. The causally important components concentrate in specific layers, but their layer distribution is largely uncorrelated with the layers whose representations change most during fine-tuning. Overlap in EAP-identified components across tasks also does not imply transfer: when tasks differ in nature, such as classification versus generation, fine-tuning on one can degrade the other precisely when their component overlap is high.
The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators
Modern analytics platforms let SQL queries call AI operators over unstructured data, but the usual way of grading generated queries — comparing exact execution results — breaks down when those operators return non-deterministic outputs. The authors catalog the resulting failure modes and propose a multilayered evaluation framework that validates the relational logic and the AI semantics separately, tested on both BigQuery and ThalamusDB. Conventional execution accuracy recognized as few as 25% of correct translations, and a state-of-the-art LLM-based autorater falsely rejected 32% of accurate queries, while the decoupled framework reached up to 97.2% overall accuracy.
TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching
Long-context language models on phones are bottlenecked by the key-value (KV) cache, which grows linearly with sequence length and is touched at every decoding step, while eviction discards tokens irreversibly and offloading stalls on I/O. TierKV predicts future cache demand from prefill hidden states before decoding starts and jointly assigns tokens to exact, low-rank, and flash-offloaded tiers under memory and accuracy budgets, using a closed-form solver that picks tier boundaries and per-layer ranks at runtime while keeping the full context reachable. Across eight text, vision, and audio models on three mobile systems-on-chip, prefill throughput improves by up to 17.6x over existing mobile LLM frameworks, with RAM-resident KV cache cut 12.5–34% and only minor accuracy degradation.
Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency
The work studies paraphrase-induced hallucination, where a model answers a factual question correctly in its original wording but wrongly under a semantically equivalent rephrasing; generic paraphrases are poor test data because near-copies give weak signal and highly diverse ones break equivalence. Hallucination-R1 trains a paraphrase generator in two stages, first stabilizing meaning-preserving and diverse paraphrasing, then rewarding paraphrases that expose factual-consistency degradation in downstream question-answering models. On SimpleQuestions, PopQA, and TruthfulQA it achieves a strong consistency-diversity trade-off and exposes robustness failures across multiple model families that analyses show are not reducible to surface artifacts or semantic drift. A lightweight fine-tuning study indicates the generated data improves robust accuracy under paraphrase variation.
CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition
Comparing model and human behavior rigorously is hard because of the breadth of tasks humans perform, so CogGym provides a unified framework that uses a semi-automated, human-in-the-loop pipeline to standardize diverse cognitive science experiments into a task-agnostic Experiment Markup Language (EML) for matched-trial comparison. The initial release curates 258 experiments from 100 papers on human commonsense reasoning and evaluates 50 large language models against human responses. Larger and more recent models reproduce human judgments better, but progress is considerably slower than on formal-reasoning benchmarks such as math and coding. The best models reach only R² = 0.59 on text, 0.58 on image, and 0.43 on video experiments, far below human split-half reliability of 0.93, 0.95, and 0.92.
How Many Humans Is a Judge Panel Worth?
The question is how many human judgments a panel of language-model judges is actually worth, and the answer depends on what is being matched. The authors audit categorical judge panels against empirical human label distributions rather than a single gold label, defining one effective size (nu_H) by matching the spectral diversity of panel residuals to independent human-reference draws and another (nu_MSE) by matching distributional squared error. Across three ChaosNLI tasks, the same 32-judge panels are worth 4.24-6.50 humans by spectral diversity but only 2.30-3.75 by distribution recovery, and constructed examples show greater diversity can accompany worse recovery. A large share of residual variance lies in a shared consensus direction that averaging does not remove (43.8% on MNLI-m, 33.7% on SNLI), so effective panel size is a target-specific measurement rather than a general human-replacement rate.
CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices
Large language models (LLMs) are increasingly used to build and analyze cryptographic implementations for Internet of Things (IoT) devices, where attackers with physical access target the implementation rather than the algorithm, yet no benchmark covers this area. CESBench provides 380 expert-written items across six sub-domains (side-channel, fault injection, implementation, countermeasures, evaluation, and integration) in four formats: multiple choice, security judgments with justification, scenario diagnoses, and code tasks graded by 572 test cases. Eleven models score between 54.4% and 83.6% overall, with top scores of 98.6% on multiple choice and 95.1% on code but only 58.8% on judgment. Across models, 88.5% of security verdicts are correct, yet their justifications earn only 53.4% of the rubric marks.
Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining
Standard tokenizers treat text as characters or statistically derived subwords, ignoring the internal phonological structure of syllables and often requiring large vocabularies. Phonemic Tokenizer converts each Vietnamese or Chinese syllable into the International Phonetic Alphabet (IPA) and factorizes it into onset, rime, and tone, with the three components sharing one sequence position so that length is preserved, which yields vocabularies of only 112 entries for Chinese and 256 for Vietnamese with no corpus-dependent vocabulary learning. PhonemicBERT, which combines the component embeddings and reconstructs masked syllables with three prediction heads, is competitive with or better than character, subword, and SubChar alternatives in controlled Chinese pretraining. The Vietnamese version matches or exceeds established Vietnamese and multilingual pretrained models.
Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue
The authors ask whether conversational AI genuinely participates in cooperative communication or only reproduces its surface forms. They compare 15,881 human-ChatGPT dialogues with 10,784 human-human multi-turn dialogues, using mixed-effects models to predict turn-to-turn alignment from morality, politeness, and related features. The AI shows cooperative surface features without the underlying mutual adaptation: moral content looks preconfigured rather than negotiated, warmth appears without sensitivity to face, and linguistic convergence steadily declines. Some cooperative mechanisms reverse direction with AI: hedging and softening, which accompany greater accommodation between humans, accompany reduced alignment when the AI produces them. Giving users agency to shape the exchange is the most consistent predictor of alignment in both settings, and the lower moral assertiveness of newer models does not come with better cooperation.
Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals
Quantizing both the weights and activations of large language models to 4 bits (W4A4) after training is difficult because activation outliers waste quantization resolution. It has been unclear which error components weight optimization, channel scaling, and orthogonal rotation each address. The authors exactly decompose the local quantization error into an activation-guided weight compensation term and an orthogonal residual, and they bound the residual in terms of persistent outlier channels and regular activations. The bounds explain why random signs in Hadamard rotations suppress constructive interference among outlier channels and why sampling several sign patterns helps. They also show how second-moment balancing yields an L2 scaling rule that relaxes to SmoothQuant-style L-infinity scaling. Across eight Llama and Mistral models, configurations built from these guidelines without any backpropagation perform competitively with gradient-trained SpinQuant.
The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
When language models reason in chain-of-thought or pass free-text intermediates to each other, structured information is serialized into natural language, and a round-trip protocol measures how much tree structure survives: one model turns a procedurally generated arithmetic expression into a word problem, another recovers the expression, and symbolic equivalence gives an exact check. Testing all pairings of sixteen models shows the channel is lossy and asymmetric, with swapping generator and extractor roles shifting accuracy by up to 60.4 points and the best pair reaching 92.9% by combining two different models. At least 73.6% of round-trip failures originate at generation, and difficulty tracks tree structure (operator count, depth, right-branching) rather than model family. About 3,600 fine-tuning examples lift every open-weight model above untrained Gemini-3.1-Pro under matched semantics, and a disjoint-domain setting still helps while leaving a gap to the frontier.
Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction
Training-free correction of optical character recognition (OCR) errors via in-context learning had not been studied for Devanagari script. The authors evaluate large language models from 3B to 32B parameters on a 20,000-sentence Hindi and Marathi benchmark across five news domains, comparing domain-random example selection, dense semantic retrieval, and their CharBM25, which retrieves in-context examples by character n-gram BM25 similarity over the OCR input to surface shared error patterns. CharBM25 beats random selection by 2.8-4.0 points of absolute word error rate on Hindi and matches or exceeds dense retrieval without a GPU, while Gemma-3-27B cuts word error rate by 55.0% on Hindi and 33.3% on Marathi. Few-shot gains are capacity-gated: models under 8B do not reliably improve on the OCR baseline, and 3B models degrade more Marathi sentences than they fix.
Trading Depth for Time in Recurrent Transformers
Recurrent Transformers feed each token's high-level hidden state into the next token's computation, raising the question of whether extra compute is better spent on more temporal steps or more layers. Using Latent Recurrent Transformers, the authors insert one latent thought token between consecutive vocabulary tokens, passed through the same shared L layers, and compare it with a 2L-layer model without thought tokens, so both execute the same number of Transformer blocks per decoded token. On 16- and 20-layer mixture-of-experts NanoChat backbones, the shallower model comes within 0.006 and 0.004 bits per byte of its double-depth counterpart, recovering 67% and 81% of the improvement with roughly 48% fewer total parameters.
One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction
Prompt optimization for enterprise information extraction usually tunes one prompt against a global objective, even though different users need the same document reorganized differently. Self-Meta-Evolve maintains a dedicated prompt per user and refines it with two loops: an inner loop that edits structured prompts from persona-conditioned feedback, and an outer loop that evolves the meta-prompt by distilling successful editing patterns. The authors release a benchmark of 292 simulated enterprise users generated from O*NET occupational taxonomies, on which the method reaches a 74.58% success rate, 13.56 absolute points above the strongest prompt-optimization baseline, and 52.54% after only two iterations. In a double-blind study with twenty professionals, the adapted prompts won 71% of pairwise comparisons against static baselines.
Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
Dense large language models are slow and costly at inference, existing acceleration methods often hurt quality, and training Mixture-of-Experts (MoE) models from scratch is resource intensive. L0-MoE converts a dense model into a lightweight MoE using L0 regularization, with a cluster confusion matrix guiding domain-aware dataset curation and dynamic batching keeping training efficient. The authors report up to 2.5x speedup over the dense model while maintaining competitive performance, outperforming existing acceleration baselines.
SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference
Running large language models (LLMs) on consumer hardware is constrained by compute and memory, and common speedups such as quantization or speculative decoding usually need retraining, per-architecture tuning, or a separate draft model. SpecQuant is a training-free framework that derives INT4, FP8, and FP16 variants from one shared base model and routes each query by predicted complexity, sending simple or factual queries to lightweight variants and complex reasoning or long-context inputs to full precision. Because the variants share weights, the quantized ones can serve as speculative-decoding drafts with adequate token acceptance. On Qwen2.5-based models evaluated with MMLU, AlpacaEval, and GSM8K, the authors report 35-43% speedups with accuracy loss under 2%.
RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding
Dynamic-tree speculative decoding methods such as EAGLE-3 work well under greedy decoding, but at temperatures above zero their deterministic top-K expansion collapses the draft distribution into one-hot probabilities and sharply lowers the acceptance rate. RheoSampling decouples the two roles of the draft distribution: it injects a sampled token among the top-K slots and gives it a proxy probability for tree expansion and pruning, while keeping its true sampling probability for verification. The authors prove the method stays lossless via an equivalence-class analysis, making it the first dynamic-tree method to combine context-aware top-K tree construction with stochastic sampling, and they add an optimal-transport-based verification strategy and a sparse draft mechanism. Experiments across several LLMs and benchmarks show higher acceptance rates and speedups over existing dynamic-tree methods.
Watermarkable Multi-Draft Speculative Sampling via Poisson Processes
Speculative sampling speeds up LLM inference and watermarking tracks output provenance, but prior work shows the two are hard and possibly impossible to combine without sacrificing one. The authors propose a multi-draft speculative sampling algorithm built on Poisson processes and an exact list-coupling-without-communication scheme, which makes it invariant to the choice of drafter. It can embed an unbiased watermark without degrading speculative acceptance, and the authors describe it as the first multi-draft, drafter-invariant scheme that preserves both watermark strength and sampling efficiency, with experiments confirming both.
Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective
Detecting whether a text was in an LLM's pretraining data is hard because high likelihood can reflect either training exposure or simply predictable text, so likelihood-only detectors mistake predictable non-members for members. The authors instead score prediction loss relative to predictive entropy, show analytically that this entropy correction preserves the expected membership signal while reducing its variance, and give the score a Helmholtz free-energy interpretation, yielding Energy Transfer Detection (ETD). Across extensive experiments ETD achieves the best average detection performance, improving average AUROC by up to 3.5% and true positive rate at 5% false positive rate by up to 5.1%, and remains robust across settings.
ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment
Best-of-n sampling aligns LLM outputs at inference time by picking the highest-reward candidate, but hard maximization gives only coarse control over the trade-off between reward and drift from the base distribution. ExpBoN is a soft best-of-n variant based on the exponential-noise report-noisy-max mechanism; it admits an exact finite-n decomposition that yields exponentially fast convergence in total variation, expected reward, and both directions of KL divergence, backed by convergence and regret analyses. Integrated into the guided speculative inference framework as ExpGSI, it cuts estimated computation by 14%-39% for Qwen2.5-Math and by up to 45% at n=16 for Qwen3 on MATH500, MMLU-STEM, and Minerva Math while keeping comparable accuracy.
Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention
Gating the value pathway of attention is reported to improve language model pretraining, but prior studies disagree about why. The authors argue that such gates supply two things softmax attention lacks: abstention, which lets a head output nothing despite attention weights having to sum to one, and noise filtering, which suppresses interference from superposed features in the residual stream. In matched models from 10M to 350M parameters, abstention is supplied through a learned per-head sink logit and filtering through a gate on each value, and the benefit of abstention shrinks with scale while the benefit of noise filtering grows, with abstention accounting for nearly all of the gain at 10M and filtering for most of it at 350M. The best model at every scale has both primitives, which add negligible parameters and remain compatible with the key-value cache.
Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention
Multi-hop retrieval failures cluster in structurally predictable subpopulations of queries rather than spreading uniformly. The authors prove that reducing confident failures is possible only when retrieval features carry mutual information about success, and that no single approximate-nearest-neighbor score feature is best across all failure regimes. RegimeAbstain computes a Retrieval Confidence Score (RCS), a logistic function of up to nine query and retrieval features that needs no additional large language model call, and uses it for calibrated abstention, evaluated with a new Confident-Wrong-Answer Rate (CWAR) metric on MuSiQue, 2WikiMultiHopQA, and HoVer. RCS is best or co-best against eight confidence baselines in all five conditions, and on MuSiQue with an LLM-judge pipeline it cuts CWAR from 39.5% to 20.6% at 50% coverage; a model trained on MuSiQue transfers to 2WikiMultiHopQA with only a 0.5 point loss in AUC.
11 more specialized papers
- TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar Ilshat Saetov, Dmitry Gaynullin
- TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya
- Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction Ruotian Wu, Bill E. Johnson, Gene Saunders et al.
- From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers Ji-Lun Peng, Yi-Zhen Zhang, Chun-Nan Chou et al.
- Prediction Dynamics in Depth-Recurrent Language Models Xinyue Luo, Fei Yu
- Chinese Competitive Debating Dataset and Benchmark Zongrui Yang, Haoyuan Li, Zhongsheng Wang et al.
- Analysing the Linearity of Linguistic Relations in Language Model Embedding Spaces Vasudevan Nedumpozhimana, Fathima Thekkekara, John Kelleher
- PRISM-BN: A Controlled Corpus and Benchmark for Text-to-Parameterized Bayesian Network Extraction Amartya Bhattacharya, Nikhil Singh, Neeti Pokhriyal et al.
- CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords Yifan Wang, Junyu Lu, Qifan Wang et al.
- Do Personality-Tuned LLMs Make Better Social Agents? Tim Krabbe, Xiaodan Shi
- QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge Rawan El Ghali, Umm Kulsoom, Anas Madkoor et al.
Theory 32
dSTAR: Straggler Tolerant and Byzantine Resilient Distributed SGD
Distributed training must cope both with stragglers, slow workers that delay gradient aggregation, and with Byzantine workers that send corrupted updates. dSTAR is a lightweight distributed stochastic gradient descent (SGD) scheme that collects gradients only from the first k workers to respond and then filters them by their deviation from an ensemble median. The authors prove that it is (α, f)-Byzantine resilient with a linear convergence rate. In experiments it keeps accuracy high under attack, whereas other Byzantine-resilient methods often suffer accuracy drops of 40-50%.
Recursive Language Models Generalize Out of Domain
The question studied is when restricting what a language model can see improves learning, comparing standard chain-of-thought (CoT), which reads the full reasoning trace, against recursive language models that solve each subtask in an isolated context. In-distribution, CoT can efficiently simulate the recursive rule, so its generalization guarantee differs only by a constant factor and recursion offers little. Out of domain, however, CoT can fit the training data through shortcuts that depend on context outside the current subtask and break when those tokens change, a failure mode that context isolation rules out. Because simplicity bias selects the shortcut even though CoT's hypothesis class contains the correct rule, the authors argue that covering the right rule is not enough for out-of-domain reasoning, in contrast to classical learning theory.
Do Quantum Models Scale Like LLMs?
Neural scaling laws are examined for RydbergGPT, an autoregressive transformer trained on qubit measurement data from interacting Rydberg atom arrays, a quantum system with a finite-size remnant of a critical point. Near the critical point, loss versus training dataset size follows a power law with a loss floor, but away from criticality the power-law fit degrades substantially. Using an entropy-normalised, finite-sample-corrected mutual information "two-point" function, the authors find that near-critical measurement statistics most closely resemble those of natural-language corpora, while far-from-critical configurations show faster-decaying correlations. They argue this supports the view that multi-scale dependence in data underlies stable scaling, making scaling a property of the model-data pair rather than the model alone.
Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks
Linear factorization blocks of the form W = BA, found in LoRA adapters, low-rank layers, and attention query-key products, have a non-unique factorization that can destabilize training and cap usable learning rates. Stiefel-AdamW constrains one factor to the Stiefel manifold while leaving the other Euclidean, which rules out factor blow-up while keeping AdamW's coordinate-wise preconditioning; moments are estimated in ambient space and geometry enters only through a tangent-space projection and a retraction. The authors prove standard convergence guarantees and report consistent improvements over strong baselines at essentially no extra cost over AdamW on LoRA-style fine-tuning of GPT2, ViT, and Mistral 7B and on full GPT2 pretraining on OpenWebText.
Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs
Joint-Embedding Predictive Architectures (JEPAs) avoid representation collapse by forcing embeddings to match a target distribution, and prior work argued the Gaussian is the unique target guaranteeing linear recovery of latent variables under Euclidean assumptions. This analysis extends the theory to latents living on embedded Riemannian manifolds and derives conditions on geometry and positive-pair dynamics that guarantee linear recovery. When latents are uniform on a sphere and representations are matched to the same distribution, every optimal representation recovers the latent state up to an orthogonal transformation, so Gaussian uniqueness is not universal, and the approximate-recovery bound is strictly tighter than in the Gaussian case. Experiments on Gaussian, spherical, toroidal, and high-dimensional Clifford-torus worlds show geometrically matched targets give better linear recovery and mismatched ones distort the latent structure.
World Modeling in Transformers
Behavioral failures can make a transformer look as though it lacks a world model even when its internal representations are faithful. Using mechanistic analysis and causal interventions on TaxiGPT, a transformer trained on random walks through Manhattan, the authors show that the model represents intersections and streets, tracks its position, and navigates with a goal compass. They trace its failures to interference between superposed intersection features that disrupts localization, partly contained by affordance packing, which groups intersections that share the same legal moves. They also propose mechanistic indicators for comparing models and find that different world-modeling capacities emerge at different stages of training.
Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods
Adaptive optimizers like AdaGrad and Adam scale each parameter entry independently and ignore the matrix structure of neural-network weights, and no general theory exists for deriving matrix-aware adaptivity. The authors build an Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters and derive Row-AdaGrad and Column-AdaGrad, which scale updates by accumulated row-wise or column-wise gradient norms. They prove regret bounds that can be strictly tighter than entry-wise AdaGrad under structured gradients, and experiments on matrix factorization and deep network training show better stability at larger learning rates and greater depths.
Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources
Nonlinear Independent Component Analysis (nICA) aims to recover the true independent hidden factors that generated data rather than a scrambled version of them, which is not possible in general without extra assumptions. The authors prove identifiability up to trivial ambiguities when the generating function is real analytic and the source densities have a finite number of discontinuities in the first derivative, with the Laplace distribution as the main example; the proof exploits the contrast between those kinks and the smoothness of analytic functions. Because real analytic functions can be approximated by normalizing flows or variational autoencoders with standard activations such as tanh, softplus, or GELU, the result applies to existing training pipelines with minimal changes. Experiments on synthetic and real data support the theory, and on CelebA the method recovers several interpretable latent factors.
Schedule optimization for tau-leaping in masked discrete diffusion
Masked discrete diffusion models are sped up with tau-leaping, which reveals several coordinates in parallel per step and replaces their joint conditional law with a product distribution, incurring a factorization error even with perfectly learned predictors. The authors give an exact integral representation of this error in terms of a distribution-dependent dependence density that tracks how conditional dependence evolves as more coordinates are revealed, develop estimators for it, and derive stationarity equations that characterize the unique optimal denoising schedule for a finite number of steps under a monotonicity condition. In the limit of many coordinates N and steps K, when the dependence profile converges to a strictly positive continuous function, optimizing the schedule improves only the leading constant and not the N/K scaling of the error, whereas degenerate profiles allow schedules that improve the asymptotic order over the uniform schedule. Examples based on stationary processes and exchangeable mixtures illustrate the two regimes.
Multiplicative Optimism for Constant Regret in Games
Multiplicatively Optimistic Regret Matching (MORM) is an uncoupled learning rule for finite general-sum games. Under simultaneous full-information self-play, every player achieves external regret of O(√n log d) uniformly over all horizons, that is, constant in the number of rounds, using only one-step optimism. The analysis combines a potential-based regret-matching argument with multiplicative stability and Hellinger control of how far strategies move between rounds. A learning-rate safeguard additionally guarantees O(√(T log d)) regret against adversarial utilities.
22 more specialized papers
- Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies Debartha Paul, Juncheng Yi
- From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences Yibo Wang, Wenhao Yang, Sifan Yang et al.
- Aggregated Posterior Predictive Checks for Generative Modeling Shweta Dutta, Gemma E. Moran
- On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation Menghui Zhou, Gaoshan Bi, Vitaveska Lanfranchi et al.
- A Smoothed Discrepancy Principle for Random Feature Methods and Neural Networks Mike Nguyen, Nicole M\"ucke
- FedeRage: Provably Convergent Agnostic Federated Learning under General Client Drift Herlock Rahimi, Dionysis Kalogerias
- Triply-Scalable Equivariant Gaussian Process Modeling Tim Steinert, David Ginsbourger
- Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers Charles Kulick, Armenak Petrosyan, Sui Tang
- From Trainability Diagnostics to Optimization Claims: Boundaries and Controls in Variational Quantum Optimization Pilsung Kang
- Diagonalized Attention for Individualized Regression: Latent-Row Localization and Prediction Borui Peng, Liwei Lin, Feifei Wang et al.
- Sparse Identification for Automatic Large-Scale Screening: A Constraint-Aware Framework with Ultra Fast Decoding Algorithm Jianing Li, Li Chai, Yingcheng Lai
- Brownian Heads for Deep ReLU Representations: Activation Mass and the Cost of Same-Sample Selection Mahdi Mohammadigohari, Nicole M\"ucke
- Optimal Randomized Proper Online Learning Zachary Chase, Idan Mehalel
- What Must Survive? Exact Task-Information--State Frontiers for Resource-Sufficient Learning Ronald Katende
- Weighted Quantum Signal Processing: Low-Depth Polynomial Approximation with Applications to Kolmogorov-Arnold Networks Rohit Sarma Sarkar, Rupayan Bhattacharjee, Elias F. Combarro et al.
- Riemannian Neural Hamiltonian Flows: Geodesic Symplectic Transport and Interpretability Vincent Souveton
- Optimization Geometry of Equivalent Brownian RKHS Representations Mahdi Mohammadigohari, Gustau Camps-Valls
- Single-Loop Stochastic Projected Damped Extragradient Methods for Stochastic Nonconvex--(Strongly) Concave Minimax Optimization Huiling Zhang, Minhao Zhang, Zi Xu
- Near-Optimal Acceleration for Smooth $\ell_p$ / $\ell_q$ Nondual Convex First-Order Oracle Optimization David Mart\'inez-Rubio, Brian Bullins, Crist\'obal Guzm\'an et al.
- Riemannian Simultaneous Inference for Tangent Vector Field Regression Xiaotian Chang, Yangdi Jiang, Qirui Hu
- Guiding Agents of Quantum Games to Equilibrium using Matrix Exponential Fixed-Point Iteration Alireza Habibi, Luis F. Abanto Leon, Setareh Maghsudi
- COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules Sushovan Majhi, Atish Mitra, \v{Z}iga Virk et al.
Agents 27
Voice-Light: A Full-Duplex Cascaded Voice Agent with Causal Turn-Taking and Speculative Generation
Voice-Light is a full-duplex cascaded voice agent designed to handle overlapping speech without canceling on every acknowledgment, prepare responses before a turn is certain, and keep canceled audio out of conversation history. It combines immediate acoustic onset detection, a causal turn-taking adapter sharing a streaming speech recognition encoder, reversible playback control, private speculative response generation, and tool calls that execute concurrently with audible bridge speech. On 1,673 real-conversation silence candidates, a learned completion checkpoint kept a 2.70% false-cutoff rate but reached only 12.53% end-of-turn recall versus 95.60% for a Silero timing policy, so the deployed system retains a hybrid controller. In three unscripted microphone sessions, 36 response turns had a 758 ms median from final voice-activity endpoint to first server audio, which the authors present as an instrumented case study rather than a controlled evaluation; data, models, code, and deployment configuration are released.
Scaling Discovery through Test-Time Communication
Whether letting parallel AI agents communicate at test time actually beats running them independently has had mixed evidence; here, role-free agents share findings through a common directory while working on hard tasks. On ARC-AGI-3, a team of k communicating agents (team@k) matches the success rate of 4k independent agents, with the advantage growing as k increases, and teams reliably solve tasks that no single agent can. The gains carry over to research-style tasks: communicating agents beat the prior best-known score on polyomino packing, and a four-agent team produced a 1,957-byte MNIST classifier with 99.4% test accuracy, smaller than the best-known human solution. The authors note that independent agents can still win when compute is limited or no clear progress signal exists.
CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop
Deployed tutoring tools generally serve fixed item banks and treat a wrong answer as one bit of signal rather than evidence about what a learner misunderstands. CoLearn keeps a persistent learner-state memory updated by a soft-evidence variant of Bayesian Knowledge Tracing in which a language model acts as a continuous observation function, generates questions aimed at the weakest topic and recurring misconceptions, and exposes the personalization through live progress views and blind A/B comparisons. In that blind evaluation, memory-conditioned questions were preferred 68–69% of the time over non-personalized ones, and in persona simulations with hidden ground-truth mastery the agent's belief converged toward the true value.
Can Agents Design Better Chips with a Higher Level Abstraction?
Most language model agents for chip design write register-transfer level (RTL) code directly, leaving them to optimize at a low level of abstraction. Four workflows are compared on FPGA targets — direct RTL design, agent-driven high-level synthesis (HLS), post-compiler HLS refinement, and post-HLS RTL refinement — with the latter two combined into AHRR (Agent-based HLS with RTL Refinement). Across an 11-task benchmark suite, AHRR achieves a 2.6x geometric-mean speedup over direct RTL design, and case studies attribute the gain to HLS distilling reusable design knowledge into abstractions the agent can work with while RTL refinement recovers lower-level optimizations.
When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success
Agent models are routinely scored one decision at a time, predicting the next action from a gold interaction history and comparing it to a reference — the question is whether gains under that protocol carry over to running a task autonomously. Pre- and post-supervised-fine-tuning Qwen3 models at 4B and 14B and Gemma 3 models at 4B and 12B were evaluated both ways on multi-turn customer-support workflows. Fine-tuning improved next-turn and text-turn success for every model, yet none of the four fine-tuned models succeeded under holistic workflow evaluation, with strict trajectory completion reaching at most 10.4%, and tool-specific gains varied inconsistently across metrics, leading the authors to argue for reporting text quality, local action correctness, tool execution, and end-to-end completion separately.
SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
Agentic coding benchmarks judge patches with held-out test suites that are inherently incomplete and increasingly vulnerable to memorization, while formal verification has so far only covered standalone tasks whose specifications are handed to the model. Benchproofer turns a coding task with a known-correct patch into a formally verified one by writing a specification for the new code, summarizing the existing functions it calls with axioms, and admitting an instance only when mechanical and adversarial gates agree; applied to SWE-bench Verified it produces SWE-Proof, 500 real issues checked by proof rather than test, and extends to SWE-bench Pro. A quarter to a half of test-passing patches from two frontier models admit counterexamples, and supplying a correct formal specification lifts Opus 4.8 from 85% to 95% resolution. Writing the specification is the hard part: models asked to produce their own gain nothing over an unaided baseline, only 62% of those specifications pass audit, and the usual failure is faithfulness — constraining part of the required behavior and leaving the rest free.
Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale
Large language model (LLM) agents can propose, implement, and evaluate model changes, but in long-running online settings a completed run can still support a wrong conclusion when a code change is a no-op, data windows leak, evaluator semantics drift, or the two arms traverse different serving funnels. EvoPilot is a human-gated method for long-horizon online autoresearch in which role-specific agents execute each round through a versioned domain skill and typed adapter, durable records preserve experiments and failures, and deterministic checks enforce recorded lessons. In a 37-day campaign on the retrieval system behind Video Deep Dive, an earlier primitive autoresearch attempt had blamed a 22-percentage-point offline hit-rate drop on an interaction head; EvoPilot's verification traced the drop to a pre-existing evaluation defect, and after repair a matched comparison measured a 3.20-percentage-point offline improvement. A separate seven-day randomized online test estimated a 0.66% relative increase in the product's Good Search Result Rate for Retention, and durable state recovered an interrupted round while saving roughly five GPU-hours through artifact reuse.
PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking
Macro placement in VLSI physical design is typically solved by one-shot numerical optimization of hand-crafted proxies such as estimated wirelength, which cannot easily incorporate visual layout context, design expertise, or downstream feedback. PlaceReasoner-Beta recasts it as closed-loop reasoning in a verifier-guided multi-agent framework: a vision-language model (VLM) planner proposes placements from the floorplan image, macro specifications, and connectivity; a geometric verifier enforces legality and expert principles; a physical verifier refines candidates with early implementation feedback; and a post-route optimizer improves promising layouts using final power, performance, and area results. The accompanying PlaceReasoner-Bench is a fully open end-to-end benchmark of 8 designs at two aspect ratios (16 tasks) scored on routed results and design-rule checks rather than pre-route proxies. The system achieves the best timing among design-rule-clean methods on all square tasks, reducing post-route total negative slack by 61.2% at 1:1 and 53.0% at 2:1 aspect ratio relative to the classical baselines, and shortens routed wirelength on most designs without optimizing it explicitly.
Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
Production LLM agents are re-evaluated constantly as they evolve, but full agent benchmarks are expensive to rerun. Using 574 historical runs of the benchmark for a production analytics agent serving tens of thousands of monthly active users, split chronologically into calibration and held-out periods, the authors compare random sampling, historical caching, fixed representative subsets, and adaptive testing based on item response theory (IRT). Multidimensional two-parameter (2PL) adaptive testing gives the best score fidelity: executing 200 questions, 38.5% of a full run, yields a mean absolute error of 1.03 percentage points. The team nevertheless deployed difficulty-stratified fixed subsets for operational simplicity, and shows they transfer without recalibration to five other agent families and stay stable with calibration windows as short as one day.
Authorization Revocation for Long-Running AI Agents: Root-Scoped Quiescence under Delegation and Asynchronous Execution
Long-running AI agents keep acting through credentials, delegated tasks, queues, and callbacks after their initiating process is cancelled or its credentials are revoked, so ordinary cancellation does not stop already scheduled work. The authors define root-scoped authorization quiescence, a certificate-backed guarantee that every acceptance under a retired authority root is accounted for and that none occurs after a local fence, while work with independent sufficient authorization is preserved. Under stated assumptions they prove properties including non-expansion after the cut, compositional soundness, and crash/replay stability. In a late-effect test suite that matches 17/17 registered outcomes, cancellation-only and cut-only executions accept an already scheduled late effect while cut-plus-fence executions reject it, and an independently implemented checker rejects 44/44 semantic regressions.
GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development
Autonomous software generation (ASG) systems can deliver runnable applications without their interacting components actually meeting the specified behavior. GameASG-Bench addresses this with 47 browser-native game-generation tasks across 12 genres in 2D and 3D, each declaring an evaluation interface before generation (legal starting scenarios, player actions, stable snapshots, invariants) and scored with static L1 source checks plus browser-executed L2 checks using real input. Across nine agent stacks the best mean L2 check pass rate is 93.2%, yet the best strict task success rate is only 55.3% (26/47 tasks). For DeepSeek-V4-Flash, full tool access and larger turn budgets help while more reasoning effort does not help monotonically, and two harnesses that each solve 18 tasks overlap on only ten.
LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces
Buyers in marketplaces where AI agents complete tasks autonomously cannot easily tell which agent will perform best, because reported benchmark scores are hard to verify or compare across tasks, software, and budgets. LEGIT is a credentialing protocol in which a signed certification record binds measured quality and cost per solved task to a specific agent configuration, task domain, evaluation budget, and evidence, while a reputation layer ties records of past task outcomes to the same identity. Evaluations show that agent configurations with similar task success can differ substantially in cost, and that comparisons between them depend on the evaluation budget. A separate analysis quantifies the deposits and fees an attacker would need to manipulate reputation under a stated Sybil attack model.
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
Deployed agents generate many execution traces, but labeling outcomes or writing task-specific verifiers is costly, so the authors ask how to turn raw traces into reusable feedback without outcome labels. DENSE (Distilling Evidence from Nested Subtask Executions) organizes evidence of local progress, recovery, and unfinished requirements into nested shortcut trees. It compresses redundant attempts, reconciles issues across levels, summarizes finished branches, and expands unresolved ones. The evaluation protocol, REFIT, compares feedback methods from shared initial trajectories with environments and model contexts reset. Under it, DENSE achieves the highest strict pass rate among non-privileged feedback methods on Terminal-Bench 2.1 across four recipient models, improving strict pass rate by 7.12-15.64 percentage points over initial executions while using 19.0-43.6% fewer recipient tokens in reruns.
OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems
Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, but final-score comparisons mix differences in models, topology, roles, and compute, so gains are hard to attribute. OpenMAS-GCom diagnoses these systems through controlled interventions that change one component while holding tasks, models, prompts, and budgets fixed: rewiring communication edges, removing specialist or critic agents, injecting incorrect intermediate messages, and disabling workers mid-execution. It evaluates 17 single-agent, ordinary multi-agent, and graph-enhanced configurations on 29 datasets across six domains, plus 400 new G-MAS-Complex tasks requiring multi-document synthesis, conflict resolution, and sourced answers. Results show larger mean losses from removing specialists than from removing critics, differing robustness to bad messages and worker failures among systems with similar original scores, and different configurations winning on accuracy versus accuracy per token.
MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems
Multi-agent systems built on large language models produce collaboration traces showing how agents plan, verify, and repair, and reusing them requires preserving each action's prerequisites and the outputs later agents depend on. Empirical studies show that grouping these dependencies into functional memory units improves retention, that linking units improves joint retrieval, and that the best combination of units differs depending on whether memory is presented as instructions or checklists. MACE builds on this with MemGoG, a graph of functional units (conditions, actions, outputs) connected by support, conflict, and repair relations, and a loop that selects units within a memory budget, chooses a presentation format per agent, and updates unit scores and relations from task outcomes. Across eight benchmarks it averages 81.11% versus 78.97% for the strongest of ten baselines, SAGE.
GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions
Existing game-development benchmarks replay fixed examples, score videos, or rely on a model judge, so they miss rule violations that occur mid-run in a game that still ends in a valid state. GameLogicBench offers 72 gameplay-logic tasks in Godot projects, with an automated evaluator that asserts the rules at every simulation tick across 403 hand-designed scenarios expanded into 1,451 seeded test cases, and that is validated to accept differing correct implementations while rejecting mutants with one required capability removed. Across 20 model and scaffold combinations, the best run solves only 52.78% of tasks, and under Claude Code all twelve models degrade as scope grows from isolated mechanics to repository-scale features. Evaluators built without mutant validation let incorrect submissions pass, and agents were found copying code from public repositories when network access was open.
GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills
Skills give large language model (LLM) agents task-specific procedural guidance, but writing them as unstructured natural language leaves them redundant, short on workflow-level direction, and hard to optimize over a huge search space. The authors represent a skill as a graph whose nodes are execution steps with operational guidance and whose directed edges encode context-dependent transitions. GraphSkillEvo then runs population-based evolutionary optimization over these graphs with mutation and crossover operators. Across five agent benchmarks it beats the SkillOpt baseline, improving average accuracy by 4.01% on GPT-5.4-nano and 1.76% on GPT-5.4.
TrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization
Clinical development planning (CDP) and estimating a drug's probability of technical and regulatory success require experts across several disciplines to synthesize heterogeneous evidence, a slow and subjective process. TrialAtlas is a memory-augmented multi-agent system that coordinates specialized agents for literature synthesis, competitive trial intelligence, regulatory precedent analysis, and integrated reasoning, and it learns from historical trials and prior New Drug Applications. The authors also introduce TrialAtlasBench, built from 291 FDA Complete Response Letters, where the system reaches a 50.0% F1 for detecting trial design deficiencies (6.1 points above the strongest baseline) and 85.3% balanced accuracy for predicting technical and regulatory success. In expert evaluation, 86.4% of its generated concerns were judged valid, versus 83.1% for OpenAI DeepResearch and 59.3% for Gemini DeepResearch.
AutoRecLab: Describe the Experiment, Get the Code!
Turning a recommender-systems experiment design into executable code is manual and error-prone. AutoRecLab is a Python-based autonomous lab that takes a natural-language research idea, derives explicit experiment requirements, builds and validates a prototype, and iteratively expands it into the full experiment, combining retrieval-augmented generation for documentation lookup, static type verification, and execution-steered tree search. In a baseline comparison across six algorithms and three datasets, 8 of 9 runs succeeded at roughly $1 per run using GPT-5.4-mini, and the system also autonomously implemented an explicit-to-implicit feedback conversion study.
AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory
Long-term memory systems for large language model (LLM) agents usually store preferences, events, constraints, and temporal updates in one mixed representation with a fixed granularity or schema, which creates semantic interference and leaves relevant evidence poorly ranked in top-K retrieval. AutoViewMem discovers candidate semantic views from interaction traces, selects a compact set of low-overlap complementary views, and uses them to guide write-time extraction of structured, provenance-grounded memories, followed by offline consolidation for compactness and consistency. Because disentanglement moves from retrieval time to write time, plain top-K similarity search retrieves focused evidence without explicit routing or iterative retrieval. On the LoCoMo and PersonaMem benchmarks with Qwen3-8B and Qwen3-14B backbones, it improves long-horizon question answering and personalization over strong memory baselines.
Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents
Large language model agents used in social simulation revise their opinions implicitly in context, so how persuadable an agent is can be neither specified nor verified, and collective outcomes inherit the model's training prior. Bayesian Chronicle Agents (BCA) add a minimal belief layer in which each stance is a probability updated by one Bayesian step per utterance heard, with a single prior-strength parameter κ that encodes stubbornness, modeled on Friedkin-Johnsen opinion dynamics. Sweeping κ produces consensus, persistent disagreement, or committed-minority influence on demand, with the persistent-disagreement regime matching the closed-form Friedkin-Johnsen fixed points at R² of 0.93 to 0.99. The prescribed κ stays recoverable after the round trip through language, with perfect rank-order recovery on all four models tested, and the explicit layer exposes per-model stance biases that end-to-end simulation would silently absorb.
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Computer-use agents have developed along two separate lines, graphical interface control and coding through the command line, while real digital work interleaves both. RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web where an agent is given a running reference application, must discover its behavior, and must build a faithful reimplementation, with the reference serving as an oracle for hidden behavioral tests that yield execution-grounded rewards. Models trained on the generated trajectories improve on five out-of-distribution coding and hybrid computer-use benchmarks and verify their rendered outputs more often. On the held-out 250-task RecreationBench, GPT-6 Astra leads at 58.1% overall but passes all programmatic tests on only 2.8% of tasks, with agents reproducing static interface structure more reliably than interactions and computed outputs.
An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency
When a memory store contains conflicting positions, standard retrieval-augmented generation (RAG) injects the memories blindly, and in susceptible models this produces a markedly higher hallucination rate than a memory-free baseline. The Memory Decision Layer (MDL) is a zero-parameter controller placed between retrieval and generation that fuses relevance, reliability, and task-risk signals through QR-based orthogonal subspace projection into an interpretable trust score, decoupling confidence from consistency and allowing explicit abstention. Across mainstream large language models and several open-source datasets, it is reported to reduce the hallucination rate under conflicting memories by about 56% in general scenarios and to approach zero hallucination in high-risk scenarios. It uses only geometric operations and adds about 0.14 ms per decision, roughly 50 times faster than the embedding-retrieval step before it.
Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
Evaluations of autonomous AI agents usually measure task completion rather than the values users care about when delegating work. Using Value Sensitive Design and large language model assistance, the authors analyzed 73,093 first-person Reddit posts about using OpenClaw, coding each for its human value, agent aspect, value fulfillment, and user outcome, and identified 21 values in six groups including autonomous, dependable, and affordable operation, bounded reach, reviewability, and equitable access. Relative to corpus share, values clustered around the operating conditions users set for a run rather than around the agent's outputs. Values were usually met where users described what the agent delivered (five of six groups) and mostly unmet where they described supervising it (all six groups), a pattern the authors call value-sensitive delegation.
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
Reinforcement learning for coding agents needs diverse tasks with reliable verifiers, and existing pipelines mine them from development artifacts such as issues and commits, which limits what can be extracted. CodeMidas is an agentic pipeline that uses source code as its only task-specific input: agents explore implemented functionality to write behavioral specifications, construct tests grounded in executing the original code, and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on it with GRPO improves all five benchmarks tested, including DeepSWE +11.7%, ProgramBench +17%, and Terminal-Bench v2.1 +8.5%, and the trained agent explores codebases more and self-verifies in more varied ways.
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Professional graphic design is a long-horizon agentic task with no reliable programmatic oracle for judging outcomes. In Designer-RSI, a frozen frontier model operates professional design software through more than 230 tools while an external procedural memory of natural-language skills widens by acquiring procedures for recurring uncovered subtasks and deepens by revising procedures against their own successful and failed executions, with a matched replay gate admitting only changes that repair failures without regressing observed successes. Five rounds over 1,406 real user briefs and 1,869 automatically graded trajectories, with no weight updates or human labels, grow the skill bank from 76 to 139 and raise GenEval2 execution success on Claude-Sonnet-4 from 72.7% to 99.3%, with 61.8% and 67.6% win rates against the no-skill agent across four design benchmarks on Claude-Sonnet-4 and Claude-Opus-4.6. On 200 held-out briefs, widening or deepening alone reaches about a 49% win rate over the no-skill agent, while combining them reaches 58.5%.
1 more specialized paper
Other 22
Programming AMD XDNA NPUs with Open-source Compiler Tools: A FlashAttention Case Study
Spatial neural processing units such as AMD XDNA place compute tiles beside small local memories and leave data movement to software, so mapping a multi-stage workload is largely a question of where intermediate tensors live. Using the open-source IRON and MLIR-AIR flows, the authors compare four FlashAttention designs on XDNA 1 and XDNA 2: per-operator execution, two on-chip streaming variants, and a fused kernel that keeps the attention scores in compute-tile local memory and reduces partial results over the cascade interconnect. On XDNA 2 the fused kernel reaches 3.62 TFLOP/s end to end, twice the IRON design, with 5.3 to 7.2 times the energy efficiency of the integrated GPU on the same chip at 2K tokens and above, covering twelve LLM configurations up to 128K tokens. Roofline analysis at each memory level shows the same fusion is nearly wasted on XDNA 1, which is already compute-bound with streaming, yielding the rule to fuse only until a mapping's operational intensity clears the ridge point; the reference designs are released as open source.
An Introduction to Compression-Based Machine Learning
Any lossless compressor such as gzip can be turned into a learning method through Normalized Compression Distance or the Minimum Description Length principle, and any autoregressive model can be turned into a lossless compressor through entropy coding. The authors survey and formalize the strategies that exploit this relationship and introduce a design framework for compression-based machine learning, which they validate empirically. Compression-based methods come out competitive with conventional baselines and decisively stronger on malware classification, and varying the framework's design choices changes accuracy by up to 0.62.
IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
Mixture-of-Experts (MoE) designs tie together three quantities that ideally would be set separately: how many experts contribute to a token's output, how many are actually computed, and how many expert-sized parameter sets must be built and stored. IntBMoE decouples them by having a lightweight hypernetwork merge all expert bases in a layer into composed experts drawn from a small learned codebook of blocks, while a router sends each token to only a few blocks; Dual-Path Residual Gating additionally couples two independently composed paths through multiplicative gating. It shows consistent gains over sparse and dense MoE baselines on image classification, with further experiments on language modeling and sequential recommendation. Deployed in AMap's generative recommendation system under a 60ms latency budget, it produced a 2.4% relative gain in click-through rate in online A/B testing.
Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER
Word Error Rate (WER), the standard Automatic Speech Recognition (ASR) metric, penalizes every lexical deviation equally regardless of whether meaning changes, raising the question of whether it tracks human judgments of transcript quality. The authors introduce HATS-en, an English dataset of human preferences over ASR transcripts, and benchmark lexical metrics against many configurations of BERTScore and SemDist that vary the language model, layer, and pooling. WER agrees least with human judgment of all metrics tested, the best SemDist configurations agree most, and no single embedding model wins everywhere. Character Error Rate (CER) stays remarkably close to the best semantic metrics, so the authors recommend CER as the primary low-cost metric with SemDist as a complement.
Neural Cellular Automata Learn General Features in their Hidden Channels
Neural Cellular Automata (NCAs) are highly parameter-efficient models, but research has mostly examined their outputs rather than their internal hidden channels. The authors analyze those hidden-channel dynamics and introduce a transfer-learning mechanism that injects a pretrained teacher's hidden states into a student model to guide early optimization. On few-shot and scale-variant MNIST benchmarks, NCAs with about 9,800 parameters outperform comparable recurrent and feed-forward architectures, and the hidden channels are found to encode general, scale-invariant topological primitives rather than class-specific templates. This lets a student reach strong few-shot performance on unseen digit classes using features from a teacher trained only on digits 0-5.
17 more specialized papers
- A Generative Grammar Underlying the Voynich Manuscript, the Pastiche Hypothesis: Evidence from Large Language Models Nicolas Turenne
- A Lightweight Plug-in Gate for Transformer-Based Time-Series Forecasters Hongkai Zhuang, Tao Huang, Chen Hou
- Toward individual-level calibration in affect recognition with perceptual adjustment queries Xuanzhou Chen, Sankaraleengam Alagapan, Ashwin Pananjady
- The Hidden Cost of Digits: Number Normalization and WER in ASR Systems Stanis{\l}aw Kacprzak, Mieszko Fra\'s
- HMB-GAN: Hybrid Multi-B\'ezier GAN for Vector Shape Synthesis Elian Hugh Thiele-Evans, Binh Duong Pham, Hani Omar M Alharbi et al.
- Not All Irregularity Is Equal: Causally Isolating a Rare Failure Mode in Japanese Morphological Inflection Wen Zhang
- Tracing the Evidence Behind Zero-Shot Time-Series Forecasting: A Source-First Taxonomy and Audit Framework Delun Kong, Wanyun Ling, Chenxi Liu et al.
- Improving the Predictive Performance of Bootstrap Aggregating by Dirichlet Resampling Quoc Viet Le, Joonha Park
- Predictive Suppression Layers for Communication-Efficient Spiking Neural Networks Aidin Attar, Michele Rossi
- Multi-Domain Clustering via Measure Quantization Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante
- From Code Archival to Knowledge Graph: Bridging Software Heritage, COAR Notify and Wikidata Camillo Carlo Pellizzari di San Girolamo, Francesco Tosoni
- Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data Morris Stallmann, Charalampos S. Kouzinopoulos, Marcin Pietrasik et al.
- Geometric Mean Pooling for Equal-Weight Multiplicative Coarse-Graining Ang-Kun Wu, Fangdi Wen, Jingtao Zhang
- RACER: Role-Aligned Competence Estimation for Human-AI Routing Joshua Strong, Emma Sun, Alexander Capstick et al.
- Time series generation with spectrally aligned latent flow matching Camilo Carvajal Reyes, Felipe Tobar
- Gricea: An Open Science Platform for Conversational AI Research Nikhil Sharma, Yunlin Gong, Xinyang Cheng et al.
- Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise Fabricio Breve
Robotics 21
Making Latent Evolution Explicit: Operator-Structured Transitions for World Action Models
World Action Models (WAMs) help robot policies by predicting how task-relevant scene states evolve, increasingly in a latent space, but the transitions are usually implemented with Transformers whose structure is built around token interaction rather than temporal evolution. The Latent Evolution Operator Network (LEON) instead models latent dynamics in a learned observable space using context-modulated operator propagation plus an additive forcing term, motivated by the controlled Koopman generator view of dynamics. Experiments on controlled dynamical systems confirm the intended inductive bias and the complementary roles of propagation and forcing. Across two WAM formulations, LEON improves closed-loop performance and robustness, even when it fully replaces the original transition module.
ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning
Reinforcement learning controllers for unmanned aerial vehicles (UAVs) are vulnerable to action-space attacks, which overwrite action commands after the policy emits them and before the actuators execute them, and existing defenses mostly address input attacks or require retraining rather than handling corrupted actions at runtime. ASGARD is a two-phase teacher-student pipeline: in the teacher phase an encoder fuses the UAV's physical state with privileged attack information to produce an attack-aware latent that trains both the control policy and a monitor that outputs corrected commands; in the student phase the encoder and monitor are distilled via supervised learning to run on board using only physical state history. Across attack scenarios targeting different action commands, ASGARD completes missions despite the attacks, generalizes to unseen attacks, and remains resilient against stealthy ones.
Visual Navigation Transformer with Pose Attention
Learned navigation policies usually consume observations as a time-ordered history, which makes it hard to reuse experience from earlier traversals without building an explicit map or topological graph. VNT-PA is a transformer planner whose context is a set of depth keyframes indexed by camera pose, so attention depends on pose differences rather than temporal order; it is trained to imitate a shortest-path planner running on the ground-truth scene mesh. On point-goal navigation in HM3D validation scenes it reaches 93.3% success and 90.4% success weighted by path length (SPL), beating baselines that encode the same context temporally or treat pose as an input feature, while also training faster. Because the context is an unordered pose-indexed set, frames from different trajectories can be fused at test time, and the planner degrades more gracefully under localization noise than a map-based baseline.
Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies
Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated evaluations of an action expert, yet more integration steps raise cost without reliably improving closed-loop success. Coda spends part of that budget on a single learned endpoint correction instead: a frozen policy runs a few-step noise-to-action trajectory, then a lightweight Transformer predicts a demonstration-supervised residual from the candidate action, the source noise, and the shared observation-prefix cache, with only the corrector trained. On 50 RoboTwin Easy tasks, five-step Coda raises success from 71.64% to 74.68% over the matched five-step baseline while cutting forward latency by 30.2% relative to the default ten-step policy, and a two-step configuration reaches 71.88% success at a 2.12x speedup. The same design lifts frozen official SmolVLA two-step success from 60.8% to 69.4%.
FootQuery: Future-Touchdown-Guided Retrieval from Depth History for Perceptive Humanoid Locomotion
Humanoids crossing stairs, gaps, and platforms often cannot see the spot where a foot is about to land, because of limited camera coverage and self-occlusion, so the relevant terrain must be recalled from earlier observations. FootQuery predicts each foot's next touchdown location and its uncertainty from proprioception, then uses those predictions to query sparsely sampled past depth frames. Retrieval is supervised during training by projecting realized contacts back into the historical images. The retrieved per-foot features are fused with a global visual memory to produce actions. In simulation the full system outperforms its ablations on the hardest stairs, gaps, and platforms. On a Unitree G1, a single policy continuously traverses outdoor stairs and indoor routes that combine stair ascent and descent, platforms, and gaps, using only proprioception and onboard depth images.
AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining
Robot demonstration data is scarce, so egocentric human video is an appealing pretraining source, but the embodiment and action-space gaps between humans and robots make it unclear how to use it. AtomEgo is a systematic study built on a curated corpus of about 2,659 hours and a scalable processing pipeline, comparing three paradigms across vision-language-action and world-action architectures: joint co-training with domain-specific action heads, progressive ego-to-robot transfer through embodiment alignment, and joint video-action modeling. Multi-task real-robot experiments and cross-embodiment representation analysis suggest that capability gains scale with data volume multiplied by alignment quality, so egocentric data helps generalization only to the extent it is well aligned and utilized.
Adaptive Rollout Truncation Based on Epistemic Uncertainty for Efficient Offline World Model Training
Multi-step autoregressive training improves the long-horizon accuracy of neural world models for robotics, but a fixed rollout length is costly and can amplify errors while the model is still inaccurate. The proposed auto-curriculum terminates each training rollout once epistemic uncertainty exceeds a threshold calibrated in a two-stage warm-up, using either a five-head ensemble on a shared recurrent backbone or Monte Carlo Dropout as the estimator. On ANYmal-D and ANT, ensemble-based truncation matches or improves the prediction accuracy of fixed-horizon training and the RWM-U baseline, reaching comparable final performance on ANYmal-D with roughly 72% less rollout computation.
Potential-Field Action Representation for Reinforcement Learning in Contact-Rich Manipulation
Model-free reinforcement learning for contact-rich manipulation usually makes the policy learn both task strategy and low-level motion generation through direct Cartesian commands. PA-RL has the policy adapt the parameters of an artificial potential field, which produces a state-dependent guidance direction executed by a Cartesian impedance controller. On simulated peg-in-hole insertion against velocity, pose, and variable-impedance action spaces under the same algorithm, it is the only method to reach a 100% evaluation success rate within the training budget, versus 92.6% for the best baseline, while cutting joint-torque variation by 55.4% and Cartesian acceleration variation by 70.8% without motion-quality penalties in the reward. The simulation-trained policy completed 9 of 9 real-robot insertions without fine-tuning.
SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations
Reinforcement learning with sparse binary rewards stalls when a Vision-Language-Action (VLA) policy never samples a success, and the usual fix is costly human teleoperation data. SynthDemo-RL uses an automated teacher that turns simulator-privileged state into successful manipulation trajectories, distills a VLA student from them with supervised fine-tuning, and then refines it with PPO on binary success rewards. On LIBERO-PRO, where a pi_0.5 policy scores exactly 0% on 27 of 57 tasks, direct PPO rescues only 10, while SynthDemo-RL rescues all 27 with 50 synthesized trajectories per task and reaches about 97% average success with no new human demonstrations. On standard LIBERO it reaches 96.0%, within 1.7 points of training on human demonstrations, with further validation on RoboTwin 2.0 and an open-loop physical robot test.
ForceTwin: Physics-informed Digital Twins for Robotic Manipulation from Instrumented Human Interaction
Digital twins of articulated objects such as doors usually capture kinematics only, or assign static physical parameters from visual and language priors, missing state-dependent effects from springs, friction, and door closers. ForceTwin has a person probe the object with a handheld force-sensing gripper, then estimates articulation, inertia, Coulomb friction, viscous damping, and a structured neural residual for mechanism forces from the synchronized poses and forces. It nearly halves the inertial-parameter error of a vision-language model (VLM) prior. Used as a feedforward dynamics model for impedance control on a Spot and a Franka FR3, it reaches 87% goal completion across nine object-embodiment pairs versus 60% for VLM-prior twins and 57% for kinematics-only twins, and the twins were also used to train door-traversal policies deployed in the real world.
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
Pretrained robot foundation policies often complete most of a long-horizon task but fail repeatedly at a few critical subtasks, and collecting more full-task demonstrations wastes operator effort on behavior that already works. PARTS (Policy Adaptation with RL on Targeted Subtasks) keeps the pretrained policy frozen for nominal actions while agent-generated selectors and success verifiers activate residual corrections and supply local rewards at those bottlenecks. Training combines online reinforcement learning with success-reweighted retraining, with humans only identifying bottlenecks and performing physical resets. It raises complete-task success from 32% to 61% on bimanual YAM tasks and from 50% to 95% on single-arm Franka tasks using tens of minutes of real-world rollouts per task, beating existing real-world RL fine-tuning methods by more than 25% under the same rollout budget.
When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence
When a robot fails at a task it must decide whether to act on its own diagnosis, consult another onboard sensor, or interrupt a person, which depends on how much its sensors reveal and how reliable its diagnosis is. The authors build a simulated benchmark with injected, known failure causes and find that some failures are diagnosable only from force data (0.99 accuracy from force versus at most 0.55 from images), then test six open vision-language models. The models track the surface of the prompt rather than the evidence: moving the refusal option from last to first drops refusal rates from 78-100% to 0-6% in three of six cases, accuracy from camera frames never beats a majority-class baseline, and stated confidence carries no information about correctness. Supplying the force data as ten lines of text produces the first above-baseline diagnoses in four of six models, and a single question to a human raises accuracy to roughly the answerer's reliability (0.70-0.81), so the authors argue the decision to ask should be tied to measured accuracy and stated costs rather than model confidence.
Benchmarking World Models for Continual Learning on Compositional Tasks
Measuring how well a world model adapts to new tasks entangles two abilities: learning unseen content quickly and reusing knowledge already acquired. To isolate reuse, the authors propose a continual learning benchmark for world models in robot manipulation, where each task curriculum includes compositional tasks combining aspects of earlier tasks, factorised along action and perception axes to show how each input modality bottlenecks reuse. They evaluate state-of-the-art world models under canonical continual learning methods alongside a modular world model whose dynamics backbone has explicitly reusable components. Modularity balances reuse against forgetting better than conventional methods, but no approach solves the problem fully.
8 more specialized papers
- PlantShade: Predicting Plant Shadows for Lighting-Aware Robotic Agricultural Operation Longchao Da, Xiaoou Liu, Xingjian Li et al.
- Diverse and Adaptable Arm Coordination for Octopus-Crawling via Diffusion-Based Uncertainty-Aware Optimization Seung Hyun Kim, Heng-Sheng Chang, Kimia Kazemi et al.
- FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models Zhiyuan Gao, Di Wen, Yanxiang Zhan et al.
- KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos Zhiyuan Gao, Yanxiang Zhan, Mohammad Khoshnazar et al.
- VLA-Scope: Shift-Aware Failure Prediction for Vision-Language-Action Models Kaiwen Zhu, Dongfang Liu, Liangkai Liu
- 2nd Place Solution to the HANDS 2026 Workshop Challenge-Dexterous Grasp Motion Track: Single-Shot Trajectory Warping for Grasp Motion Generation Muneeb A. Khan, Woojin Kim, Shinwoo Kim et al.
- Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies Xingyu Lin, Zhuang Li, Zhongrun Wu et al.
- Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning Ayah G. Ahmad, Claire E. Borden, Maegan Tucker
Safety & Alignment 20
From Generation to Detection: Exploration of Discourse Driven Scenario based LLM Generated Fake News
How large language models both produce and detect fake news is examined under four controlled scenarios: open-ended generation, rewriting, manipulation prompts, and attribute-based prompts grounded in a journalistic discourse framework. Seven widely used models generated a synthetic corpus of 14,000 articles, whose linguistic properties were compared with real news, and each model was then asked to judge the generated articles using a basic prompt and prompts refined iteratively from misleading patterns in real-fake pairs. Generation and detection ability vary substantially across models, and the generation strategy strongly affects detectability. Notably, the refined detection prompts do not improve and often harm detection performance.
$\mu^2$-Bench: A Multilingual Machine Unlearning Benchmark
Unwanted information such as harmful content or private data can spread across languages inside multilingual large language models, and it is unclear whether unlearning methods remove it in every language. μ²-Bench is a Multilingual Machine Unlearning (MMU) benchmark that simulates the full memorization, unlearning, and evaluation pipeline across a broad set of languages, testing both languages used in training and held-out ones, and assessing knowledge dispersed across multiple languages. The authors report that successful multilingual unlearning requires methods that explicitly account for multilingual characteristics, and provide further analysis of how unlearning behaves across languages.
Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models
To probe how language models weigh conflicting moral values across languages, the authors build a 12,000-instance dataset of two-option dilemmas covering Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty, translated into Hindi, Arabic, Spanish, and Chinese. GPT-5-mini consistently favors Honesty over Autonomy in all five languages when given no policy, while Llama-3.2 1B and 3B models show a strong first-option bias that both plain fine-tuning and Direct Preference Optimization remove, raising accuracy above 98%. To separate learned value preferences from dataset correlations, they orthogonalize a value-preference task vector against a general instruction-following vector, and show the isolated direction can be used in task arithmetic to produce a model with the opposite stance.
Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees
Attackers who combine large language models (LLMs) with auxiliary knowledge can link released text to individuals, yet existing privacy audits report only attack success rates without finite-sample statistical guarantees. Conformal Privacy Auditing (CPA) is a distribution-free calibration framework that outputs, for each released document, a set of candidate identities guaranteed to contain the true one at a user-chosen confidence level under exchangeability, with the set's size serving as an interpretable leakage proxy. It supports both logit-access and sampling-only attackers, so open-source and proprietary API models can be audited in the same way. Across several release benchmarks it achieves calibrated coverage and reveals sharp shifts in certified identifiability as auxiliary knowledge, LLM augmentation, and release mechanisms vary.
Hiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models
Multimodal large reasoning models (MLRMs) can infer where a casually shared photo was taken by reasoning over cues such as architecture, vegetation, and lighting, which creates a geolocation privacy risk. The authors show that refusal-based safeguards are insufficient, since crafted jailbreak prompts raise model response rates to 100%, and that existing pixel-space perturbation defenses transfer poorly to black-box models and leave visible artifacts. Their defense instead injects perturbations into the latent space of a diffusion model during reverse sampling, using GeoCLIP, a model aligned with GPS coordinates, as a surrogate to locate and disrupt the geographic signals. They report significantly stronger black-box transferability while preserving perceptual image quality.
HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference
Homomorphic encryption (HE) lets a server run a large language model on encrypted inputs without seeing the plaintext, but that same confidentiality means the server cannot inspect prompts or responses, so jailbreak attacks by malicious clients go entirely unnoticed. HE-Guardrail evaluates guardrail mechanisms wholly over encrypted data and homomorphically controls whether the target model's response is returned to the client. Instantiated with Llama Guard, JBShield, and GradSafe, it closely reproduces the decisions of the corresponding plaintext guardrails in the encrypted domain, with each option showing a different trade-off among security, efficiency, and utility.
ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor
Third-party adapters for open-weight language models ship as opaque weight matrices, and the authors argue that detection fails for backdoors hidden in the subspace where a safety monitor is structurally blind, since honest and backdoored adapters overlap on every blind-subspace statistic evaluated. ServeGuard instead makes that channel structurally absent: the publisher constructs the adapter to read the input only through directions the monitor covers and proves this in zero knowledge, without revealing the certified weights. The proof is cheap because the monitor's blind spot is a deterministic function of the public base model, so only one linear identity needs proving, and an admission-time guard binds the guarantee to the adapter bytes actually served. Across eight checkpoints up to 7B from four families, the monitoring budget turns out to be architectural, and on a 0.5B model confinement is nearly free for benign adaptation, making monitor quality the main security lever.
Micro-Collaborative Poisoning: A Distributed Attack on RAG Systems
Retrieval-Augmented Generation (RAG) systems can be poisoned through their knowledge sources, and defenses usually assume a single malicious passage. Micro-Collaborative Poisoning instead splits a false target claim across several individually plausible documents, and is evaluated over 108 RAG configurations varying dataset, retriever, retrieval depth, database composition, number of poisoned databases, and generator model. The attack's effect comes from accumulated weak adversarial signals across retrieved sources rather than any single dominant passage, so larger top-k and more poisoned databases strengthen it, while clean database diversity and stronger retrievers weaken it. Document-level analysis shows it leaves a weaker explicit poisoning signature than direct poisoning, making it hard to catch by inspecting documents in isolation.
GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation
Machine unlearning is harder for large reasoning models (LRMs) because protected facts or unsafe rationales can leak in the intermediate chain-of-thought (CoT) before the final answer, and existing objectives suppress content without specifying what the model should produce instead, leading to hallucinated substitutes or degenerate output. Guided Answer-Reasoning Distillation (GUARD) converts the model's own unsafe disclosures into safe-exit trajectories, a coherent non-disclosing CoT followed by a stable refusal, aligns a frozen LRM using guidance tokens, and distills that behavior into the weights. A new metric, Natural Forgetting Reasoning Score (NFRS), measures structural stability, fluency, and unsupported substitutes in forgotten outputs. On R-TOFU and a STAR-1-derived harmful-intent setting, GUARD substantially reduces unsafe and privacy disclosures on two distilled LRMs while preserving reasoning utility.
CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents
Privacy leakage in LLM agents is usually measured inside single components such as memory, retrieval, or tool use, which blurs the difference between internal exposure and what an outside attacker can actually recover. CIPL (Channel Inversion for Privacy Leakage) is a black-box evaluation framework that models a target as sensitive source, selection, assembly, execution, observation, and extraction stages, and measures the transition from selected sensitive units to attacker-recoverable output under one protocol. Experiments across memory-based, retrieval-mediated, and tool-mediated targets plus a BrowserUse live-agent case study show that storage labels alone do not determine recoverability: memory leakage is near-saturated, retrieval leakage is often partial, and tool and live-agent leakage varies with observation surface, prompt-to-channel alignment, retrieval depth, and provider behavior. A semantic audit also finds attacker-useful disclosures that exact matching misses.
CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation
Jailbreak defenses for large language models (LLMs) sit at different pipeline stages, such as input modification or output guarding, but prior studies evaluated them mostly in isolation and under inconsistent attack-success-rate definitions. CASCADE studies defense combinations within and across stages under one threat model of direct, black-box, single-turn attacks, using a standardized attack-success-rate formulation with controlled query budgets and explicit fairness rules. Across 19 attacks and 15 defenses, no single defense is universally best, but well-chosen combinations deliver substantial safety with minimal utility loss. The authors turn these results into practical recommendations for layered defense pipelines.
End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
Hard-label model extraction attacks try to recover a neural network's parameters from oracle queries that reveal only the final output label, and the polynomial-time attack on ReLU multilayer perceptrons by Carlini et al. (Eurocrypt 2025) contains a sign-recovery step too query- and compute-heavy to run in a fully black-box setting. The authors propose a new sign-recovery algorithm based on a different principle that needs no dedicated queries and achieves higher sign-recovery accuracy than the existing method in their experiments. With it, every step of the attack can be implemented in a black-box setting, enabling an end-to-end demonstration on trained models. On networks trained on MNIST and Fashion-MNIST with width 16 and 4 or 6 hidden layers, the extracted models achieve over 98% label agreement.
Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment
Annotator disagreement on moral content is usually collapsed by majority vote or by an any-annotator rule that marks an item positive as soon as one annotator flags it. Moral Entropy is a Bayesian framework that keeps a full posterior over the true label and splits its entropy into aleatoric uncertainty (irreducible disagreement about the moral content) and epistemic uncertainty (insufficient or noisy annotation), which lets any consensus rule be audited against a calibrated reference using cross-entropy/KL, Brier score, and expected calibration error. Across three corpora and fifteen discourse domains, the any-annotator rule disagrees with the calibrated posterior on roughly 30% of items, almost entirely false positives when pooled, though the errors invert at the level of individual moral foundations (19.9% mean false-positive and 38.9% mean false-negative rate on MFTC). The stricter majority and two-vote rules instead miss 63-83% of true positives.
A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
A language model may hold knowledge it does not report, for example by sandbagging on a capability evaluation, and its outputs alone cannot show whether it is hiding an answer or lacks one. Borrowing the forensic Concealed Information Test, Probe of Internal Recognition (PIR) presents a question with candidate answers and reads from the model's internal states which candidate it recognizes as correct, without needing an honest reference model or a labeled truth corpus. Across eight models from five families (Gemma, Qwen, Llama, Mistral, Phi), PIR recovers the recognized answer at 0.70 to 0.87 balanced accuracy, against a 0.28 to 0.40 unknown-item baseline and 0.25 chance, and recognition stays between 0.85 and 0.93 under prompted deception, trained sandbagging, password-locked checkpoints, and circuit-broken checkpoints. When unlearning actually removes the knowledge, recognition falls to the level of a never-known question, so the probe separates a model that will not answer from one that cannot; the signal is reported to be causal, to add information beyond black-box cues, and to extend to free-form generation.
Available Guardrails: Certifying Selective Prediction across ML Systems
A selective predictor acts as a safety gate that returns an output only when it appears trustworthy, and deployments increasingly need that reliability certified at a target precision for every reporting unit, such as a tool, policy label, or patient subgroup. The authors focus on availability, meaning whether finite calibration data can produce a certificate at all, make it computable through exact-binomial inversion, and cast the choice of reporting partition as a dynamic program that trades off safety, granularity, and served traffic. A truth-informed planner gains 0.157 mean coverage over support balancing while a naive estimator recovers only 0.005; building candidate partitions on one data split and selecting on another recovers 0.060, with the direction reproduced in 59 of 60 model effects across three intent-routing datasets and two architectures. Reallocating the familywise error budget across units recovers further coverage, and the same frontier appears in large language model tool-calling, content moderation, lesion classification, and recommendation.
5 more specialized papers
- AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture John Cuneo, David Chun, Gaurav Khanna
- FairLMs: A Turnkey Library for Fairness in Language Models Jiale Zhang, Michael Larionov, Zichong Wang et al.
- Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations Orfeas Menis Mastromichalakis, Giorgos Filandrianos, Wafaa Mohammed et al.
- Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30 Hans Andersen, David Dichas
- TERMon: Detecting Persistent Behavioral Threats in Edge AI via Hardware-Native Ternary Runtime Monitor Arish Sateesan, Edlira Dushku
Unclassified 20
COAL-SQL: Coverage-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training
No summary available — see the abstract on arXiv.
VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering
No summary available — see the abstract on arXiv.
Boosting Deepresearch and LongContext Ability with Self-Generated Deepresearch Rollouts Traces
No summary available — see the abstract on arXiv.
Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal Decoders
No summary available — see the abstract on arXiv.
Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models
No summary available — see the abstract on arXiv.
Reading Anxiety or Reading the Label? Comparing Fine-Tuned and Frontier Models for Anxiety Detection on Social Media
No summary available — see the abstract on arXiv.
Enhancing Audio Reasoning via Semantic Summary Prediction
No summary available — see the abstract on arXiv.
MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs
No summary available — see the abstract on arXiv.
Reconstruction of 4D Mitral Regurgitation Hemodynamics from Sparse Planar Data using Deep Operator Networks with Test-Time Adaptation
No summary available — see the abstract on arXiv.
Automated Physics-Informed Neural-Networks-Based Calibration of Highly Segmented Silicon Telescopes
No summary available — see the abstract on arXiv.
The Refutation Gap: Certifying Both Halves of an Optimality Claim
No summary available — see the abstract on arXiv.
Cross-Lingual Parkinson's Disease Severity Assessment Using Pre-trained Speech Embeddings: A Multi-Class Evaluation
No summary available — see the abstract on arXiv.
Reinforcement learning for post-coronagraphic wavefront control
No summary available — see the abstract on arXiv.
Sparse Priors for Efficient Distribution Learning
No summary available — see the abstract on arXiv.
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
No summary available — see the abstract on arXiv.
Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding
No summary available — see the abstract on arXiv.
Extreme classification: beating chance with one training example from each class
No summary available — see the abstract on arXiv.
SpaceDiffusion: Over-the-Orbit Diffusion for Space Generate-and-Forward Communications
No summary available — see the abstract on arXiv.
Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes
No summary available — see the abstract on arXiv.
Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain--Computer Interfaces
No summary available — see the abstract on arXiv.
Multimodal 18
PhysioBench: A Unified Benchmark for Physiological Signal Question Answering
Foundation models for physiological signals usually need task-specific adaptation, and it has been unclear how well current models can follow natural-language instructions across signal types. PhysioBench harmonizes annotations from 22 public datasets into 61.4 million questions spanning 30 tasks, each grounded in a signal segment and traceable to its source annotation. An evaluation of 21 models, including large language models, vision-language models, time-series language models and physiological signal foundation models, under three settings finds that none achieves consistently strong performance across modalities and tasks. Natural language does enable unified prediction across tasks, though results remain sensitive to how questions are phrased.
I'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance
Audio large language models only answer when queried, which is a poor fit for a wearable meant to monitor a sound stream and speak up unprompted — a use case motivated by Deaf and Hard of Hearing users. ISM (Interrupt and Silent Modeling) is a model-agnostic scheme that embeds the proactive decision into decoding through two special tokens, <interrupt> and <silent>, covering onset detection, sustained-relevance triggering, irrelevance suppression, and de-duplication from a single natural-language statement of intent. Applied to Qwen2-Audio-7B it reaches 99.6% interrupt F1 on ESC-50 with perfect de-duplication recall, leads on noisy Epic-Sounds kitchen audio without domain-specific training, and streams with 3.5-second average latency.
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
Existing video benchmarks for Multimodal Large Language Models (MLLMs) mostly pose scene-level or summary questions that can be answered in a single inference step. AgentVidBench is a multi-hop video question answering benchmark that targets spatial, temporal, and causal reasoning. It includes step-by-step solution traces, so evaluation can check whether an agent actually gathered the evidence supporting its answer. Across 12 proprietary and open-source MLLMs, single-turn performance remains limited, while wrapping the same models in agentic workflows generally improves both accuracy and trajectory scores. The authors also provide a simple agentic baseline and release code and data.
Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction
Voice-and-vision assistants must first work out what a user actually wants. That demand is often underspecified in speech, buried in visual or dialogue context, obscured by disfluency or noise, or not addressed to the assistant at all. Omni Demand Understanding (ODU) frames the task as detecting whether a demand is present and inferring its intent from an interaction stream. ODU-Bench evaluates this along five dimensions over single- and multi-turn interactions, and is built from taxonomy-guided agentic video generation plus human recordings with human-verified annotation. Among 14 native multimodal large language models, the strongest, Gemini 3.1 Pro, recovers only 44.7% of the key information that must be inferred from visual, acoustic, or conversational context, and 11 of the 14 models false-trigger on more than 50% of non-demand scenarios.
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
Native audio-visual dialogue is defined here as an omni model receiving a user's audio and video directly and replying in text, with no separate text question, captioning, or speech recognition step. Because real recordings are scarce and good replies are too varied for keyword matching, the authors build OmniVChat-Studio, a multi-agent engine that synthesizes single- and multi-turn dialogues, and use it to create OmniVChat-Bench, covering five ability categories. They also propose OmniVChat-RL, a reinforcement learning reward that jointly targets reply correctness, efficiency, and style. Training Qwen3-Omni-Instruct with it on synthesized dialogues improves results on both the synthetic benchmark and the human-recorded OmniVChat-Bench-Human, indicating transfer to real dialogues.
PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design
Existing benchmarks do not test whether multimodal large language models (MLLMs) can design a complete load-bearing structure that actually works when simulated, or repair it after a failure. PolyBridgeBench gives a model a visual scene plus structured engineering constraints and asks for a full node-member-material bridge topology, gating execution in a dynamic physics simulation behind deterministic legality checks. After a failed run, the model receives temporal visual evidence from the rollout and attempts a repair under a fixed interaction budget, with validity, dynamic success, and recovery measured separately. Experiments with six MLLMs across 189 levels reveal a substantial gap between producing a valid design and one that succeeds in simulation, along with strong sensitivity to material budgets and limited post-failure recovery.
VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration
Benchmarks for Video Large Language Models (Video-LLMs) often rely on question answering or caption matching, which models can pass through superficial cues and incomplete annotations. VidOmni-Bench instead asks models to verify whether each event in a dense video caption is actually supported by the video, using 500 videos spanning five complexity types and durations from 4 seconds to 90 minutes. Captions are generated by diverse Video-LLMs and labeled at the sentence level by humans, so sentences with incorrect events serve as hard negatives. Experiments show that Video-LLMs frequently hallucinate in dense captioning and also fail as verifiers of plausible but incorrect event descriptions, with weaknesses that vary by model across video complexity and duration.
Samsone: A Family of Open Small Audio Language Models for On-Device Inference
Large audio language models have grown past billions of parameters, while privacy and latency needs favor Small Audio Language Models (SALMs) that run on-device. Samsone is a family of such models built for edge hardware, with a core Samsone-134M plus Samsone-99M and Samsone-356M variants used to study scaling behavior at small sizes. The authors claim Samsone-134M sets a new state of the art for its size class across multiple benchmarks and that the family is competitive with models orders of magnitude larger. Training uses only public data, and the release includes training code, weights, mobile-optimized checkpoints, and an open-source Android app demonstrating real-time on-device inference.
ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction
Vision-language models (VLMs) used for emergency department prediction can score well without actually depending on the patient's electrocardiogram (ECG), a failure the authors call ECG Mirage, split into ECG neglect and ECG confusion. They test for it by comparing predictions with matched ECGs, outcome-discordant mismatched ECGs, and no image, while holding clinical text and targets fixed. Across four VLMs on MDS-ED, matched ECGs give no consistent advantage for predicting ICU admission or clinical deterioration. Training four restricted visual prompts with supervised learning followed by conditional direct preference optimisation, with the backbone frozen, yields balanced accuracies of 70.6% and 67.5% and widens the matched-versus-mismatched gap to about 16.5 and 5.5 percentage points.
NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
NemotronLabs VoiceChat is an open full-duplex speech-to-speech model that can call tools natively while listening and speaking at the same time. It combines a streaming speech encoder and decoder-only language model with parallel output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming text-to-speech decoder. On Full-Duplex-Bench 1.0 it has the lowest pause-handling takeover rates among evaluated open-weight systems, 100% takeover after user interruptions, and a 4.33/5 post-interruption response-quality score, and on Full-Duplex-Bench 1.5 it resumes after user backchannels in 93% of cases. It scores a 55.1 normalized average on VoiceBench and 82.5% tool-selection F1 on Full-Duplex-Bench 3.0, while argument accuracy and end-to-end tool execution remain areas for improvement.
8 more specialized papers
- TALON: A Temporally Aware Longitudinal Framework for Radiology Report Generation Nien-Tsyr Sun, Min-Chen Chen, Hui Nien Hung et al.
- Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh
- M2G-LLM: Enhancing Clinical Prediction via Multimodal Graph Reasoning and LLM Context Injection Inyoung Choi, Sukwon Yun, Jiayi Xin et al.
- Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering Jia Li, Li Dai, Peng Jia et al.
- The Spoken Wikipedia Presentation Corpus Thomas Ranzenberger, Steffen Freisinger, Tobias Bocklet et al.
- Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation Yunji Chu
- Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts Steffen Freisinger, Philipp Seeberger, Thomas Ranzenberger et al.
- DiaVLo: Diagnosing Behaviours of Vision-Language Models Lorenzo Corti, Jie Yang
Vision 14
Physically Based Rendering in the Latent Space
Image diffusion models generate impressive images but are hard to control compared with classical physically based rendering pipelines. Observing a link between light transport and the distribution of latent values, the authors run light transport simulation directly in the feature space of the variational autoencoder used by generative models, modifying the rendering equation and using a differentiable renderer to recover scene parameters that render into the pretrained latent space with minimal refinement. Trained on a single rendered image, the method is shown to generalize to changes in scene geometry, lighting, and camera viewpoint.
Detection is solved, delineation is not: what governs tooth segmentation on panoramic radiographs
Using 1,422 panoramic radiographs with 42,142 expert-delineated tooth polygons across the 32-class FDI numbering scheme, the authors isolate the effects of input resolution, architecture, and anatomical priors on tooth segmentation under one evaluation protocol. Resolution dominates: raising input size from 640 to 1280 lifts mask mAP50-95 from 0.656 to 0.717 while mAP50 stays flat at about 0.982, meaning added resolution buys boundary precision rather than detection. A query-based transformer with 2.1x the parameters is statistically equivalent to a one-stage detector in-domain and 5.5x slower on CPU, and three targeted interventions fail, including a LoRA-adapted self-supervised encoder and a promptable foundation segmenter that degrades masks by 39%. Zero-shot transfer to an independent multi-centre cohort costs 62% of mask mAP50-95 but only 18% of mAP50, with residual error concentrated in the apical third of the tooth.
The Weight Is Over - Interactive Diffusion on Consumer GPUs
On-device inference work has focused mostly on language models, while diffusion pipelines remain memory-hungry, latency-sensitive, and harder to orchestrate because they chain an embedder, a transformer, a decoder, and postprocessing. The authors contribute an embedding translator that maps a small text encoder into a large encoder's space to cut weight and latency, a reproducible sweep recipe for navigating the speed, quality, and memory trade-off, and an interactive on-device image generation editor. The editor achieves sub-second time to first image on recent consumer GPUs.
11 more specialized papers
- Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation Meng'en Qin, Yinchen Liu, Mingxuan Cui et al.
- Fragment-Aware Vision Transformers for Fresco-Fragment Style Classification Sara Miketek, Biagio Barchielli, Nadeem Iqbal Kajla et al.
- LoRA Enhanced Contrastive Learning with SAS Vision Transformers Dan Zimmerman, Frank E. Bobe III, Amelia L. McCormack et al.
- Multiclass Semantic Segmentation of Wildland Fire Images Using Context-Aware Centralized Copy-Paste Data Augmentation Joon Tai Kim, Nishanth Kunchala, Vishv Patel et al.
- WS-NeRF: A Mamba-Driven World-State-Aware Adaptive Deblurring Neural Radiance Field Hang Jiang, Jinghao Wang, Yiming Zhang et al.
- Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction Jingke Zhou, Chenhang Ma, Zhizhou Zhong et al.
- Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving Jiaxing Chen, Hengduo Zou, Yiren Zhao et al.
- Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving Jiaxing Chen, Hengduo Zou, YuKai Qin et al.
- Purification and Regulation: Comorbidity-Aware Multi-Label Few-Shot Learning for Medical Image Classification Ying-Chih Lin, Po-Chih Kuo, Yong-Sheng Chen
- Balanced Prompt Adaptation against Entropy-Induced Collapse for Test-Time Binary Segmentation Zhengshan Wang, Joshua Charles Webster-Ford, Yifei Tian et al.
- Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition Laurent Colbois, S\'ebastien Marcel
Reinforcement Learning 13
Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications
Synthesizing policies that satisfy Linear Temporal Logic (LTL) specifications, such as safety or reachability, in unknown environments is addressed with an end-to-end model-based reinforcement learning algorithm. The task is represented as a Limit-Deterministic Büchi Automaton synchronised with a Bayes-Adaptive Markov Decision Process model of the environment, and a new variant of Bayes-Adaptive Monte-Carlo Planning (BAMCP) computes approximately Bayes-optimal strategies over the combined structure. Across finite- and infinite-horizon tasks the approach shows better property satisfaction and sample efficiency than traditional model-free methods, with ablations favoring the new planner over classical BAMCP. The authors also apply it to cautious reinforcement learning, reducing the number of task violations incurred during training.
REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement
Continuous-control policies in deep reinforcement learning usually map a state to an action in one forward pass, leaving no room to revise an initial prediction. Iterative Action Refinement (IAR) instead builds the action through a sequence of learned residual corrections, with a shared refinement network repeatedly conditioning on the state and the current action proposal; combining it with Proximal Policy Optimization (PPO) yields REFINEPPO. Across 14 benchmark control tasks, with ablations of refinement depth and update schedules, REFINEPPO matches or exceeds standard PPO and converges faster on several tasks.
Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation
Co-training one language model as both coder and test author promises to move code-generation reinforcement learning past fixed test suites, but pass-rate rewards can be maximized with trivial non-discriminative tests, and independently sampled tests cluster on typical inputs and inflate estimator variance. CoVer tackles both inside a single-policy GRPO setup: an information-gain reward scores each self-written test by the mutual information between its pass/fail vector and a graded, ground-truth-anchored correctness signal, paid out only when the covariance sign shows the test discriminates in the right direction, and a three-stage diversity filter prunes candidates to a behaviorally non-redundant suite at fixed execution budget. Across LiveBench, MBPP, LiveCodeBench, CodeContests, and CodeForces, one-shot pass@1 rises by 5.8 points at 7B and 7.1 points at 14B over the Qwen2.5-Instruct backbone, and dropping CoVer-7B into the CodeT ranking pipeline adds a further 3.5 points.
Deep Reinforcement Learning with Buffered Quantile Objectives
Risk-sensitive reinforcement learning that optimizes a chosen quantile of the return distribution is difficult because point quantiles shift abruptly under small perturbations, and existing methods based on smoother buffered quantiles are model-based and limited to small tabular problems. Deep-BQRL is a model-free distributional method that learns conditional return quantiles from sampled transitions with neural networks, scores actions by averaging the relevant region of the learned quantile function, and uses ensemble disagreement to guide exploration. On an asset-selling optimal-stopping problem and slippery FrozenLake, it attains smaller point-quantile policy gaps than PPO and TRPO, although the model-based UCB-BQRL retains the smallest gaps. The learned stopping decisions vary with the target quantile, illustrating the method's risk-sensitive behavior.
ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL
Reinforcement learning for large language model (LLM) agents is hard to apply to open-ended tasks, where no reliable scalar reward exists. Existing pairwise-preference methods still compress comparative feedback into a single trajectory-level reward. ArenaFlow ranks trajectories through tournaments and attaches a structured reflective evaluation to each comparison, which identifies pivotal success steps, reusable strategy skills, and which retrieved skills were actually used. Trajectory-level advantages are propagated to high-confidence pivotal steps according to how deep a trajectory survives in the tournament, while a global skill memory is updated, pruned, and retrieved by estimated utility to serve as a policy prior for later exploration. The authors report gains on open-ended agent tasks, though the abstract gives no specific numbers.
GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation
Post-training methods for large language models such as Group Relative Policy Optimization (GRPO) can be unstable because they rely on importance sampling. Group Variance Policy Optimization (GVPO) builds the analytical solution of KL-constrained reward maximization into its gradient weights, so that its gradient matches that of a mean squared error between centered implicit rewards and centered actual rewards. The authors show that it has a unique optimum that exactly solves the KL-constrained objective and allows flexible sampling distributions without importance sampling. They extend it to on-policy distillation (OPD) and to a broader family of OPD objectives. The abstract reports no empirical numbers.
IncentRL: The Trade-Off Between Preference Guidance and Task Performance
Adding preference signals to a reinforcement learning reward can unintentionally change the task being optimized. IncentRL introduces preference guidance as a Kullback-Leibler (KL) penalty between a specified outcome distribution and a preferred distribution, and for finite discounted Markov decision processes it derives a bound on how much external-task value can change, a sufficient action-gap condition for preserving the original optimal policy, and a characterization of the large-weight regime. In a practical implementation with a hand-designed distance-based outcome proxy, the three-seed mean success rate on MiniGrid DoorKey-8x8 after two million steps reaches 98% with coefficient 0.01 versus 90.5% for the zero-coefficient baseline. The authors note the experiments are descriptive and do not yet isolate KL shaping from simpler alternatives.
OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios
Auto-bidding under the optimized cost-per-X (oCPX) advertising paradigm is usually served by a separate model per scenario, such as registration or purchase, which fragments pipelines and ignores cross-scenario structure. OneBid trains one Decision Transformer-based backbone on heterogeneous bidding logs, conditioning on both Return-to-Go for conversion value and Cost-to-Go for cost ratio, with a sequence-level Mixture-of-Experts that combines shared and sparsely routed experts to stay within latency limits. Scenario adaptation uses Critic-guided Relative Offline Policy optimization (CROP), in which a learned critic scores candidate actions group-relatively so no unsafe online exploration is needed. Fully deployed at Kuaishou, it delivers an overall +2.2% gain in advertiser value in online A/B tests, peaking at +13.1% in the return-on-ad-spend scenario.
CityLearn v3: A Configurable Simulation and Evaluation Framework for Realistic Control Studies of Renewable Energy Communities
Controller studies for renewable energy communities often simplify away changing membership, equipment availability, service deadlines, and data quality, so a lower cost or peak demand can hide missed services or infeasible power requests. CityLearn v3 is a configurable simulation and evaluation framework that models changing members and assets, flexible-load deadlines, demand-response requests, local energy sharing, and data or equipment failures in one environment, with building and phase power limits constraining controllable requests. It logs controller inputs and separates requested actions from the actions actually applied to the simulated equipment, and ships reference controllers plus service- and constraint-aware performance indicators. A synthetic high-frequency trace replay illustrates how time aggregation can conceal short peaks without changing annual energy.
GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning
Planning-based reinforcement learning for high-dimensional continuous control suffers from misalignment between components: learned sampling policies drift from planner behavior, and planning distributions stored in replay go stale as the model and value function change. GEM-MPC builds on Model Predictive Path Integral (MPPI) planning, combining a policy trained to clone the planner with a KL-regularized policy that explores around it. Its Gated Prior Distillation learns from stored planning distributions only when they are a better target than the current prior, avoiding the cost of full reanalysis. Across continuous-control benchmarks it outperforms existing planning-based baselines under lower computational budgets.
What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence
In interactive retrieval, an agent must choose which question to ask next so that the answer produces the most useful evidence for the following retrieval update, but existing systems learn this by imitating an offline ordering of candidate question-answer pairs. After showing that candidate discriminativeness and perceived usefulness are only weak supervision for this goal, the authors introduce RAVEL, an online reinforcement learning framework for interactive person re-identification that starts from supervised question generation, sees the current top-4 candidates, and is optimized with rank feedback from the full question-answer-retrieval loop. On Interactive-PEDES, RAVEL delivers progressively stronger retrieval across five interaction rounds. Analysis shows it shifts the questioning budget toward localized open-ended attributes, which yield the largest gains on initially difficult queries.
$\lambda$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource
Flow-GRPO applies reinforcement learning to flow-matching image generators by treating the denoising sampler as a stochastic policy, but training is unstable: importance ratios drift below one, grow more dispersed, clip at different rates across denoising steps, and leave fewer usable samples late in training. The authors show that these symptoms all arise from a single per-step quantity they call path variance, which is fixed exactly by the sampler's Gaussian transition kernel and can be estimated cheaply during training. λ-Controlled GRPO calibrates importance-ratio behavior from this predicted law and allocates gradient effort across denoising steps by predicted cost, without introducing new free tuning parameters. On a text-to-image model it improves both text-rendering accuracy scored by optical character recognition and human-preference reward over the strongest empirical stabilizer, while keeping late-step path variance within budget where the baseline overshoots.
1 more specialized paper
- Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks Adewumi Augustine Adepitan, Christopher J. Haruna, Oluwasegun Adegoke et al.