Monday, September 7, 2026
Highlights
Iris: Climbing to the Search Frontier
Iris-mini and Iris-pro are search agents trained at the 35B-A3B and 397B-A17B scales, released together with their data pipeline and training recipe. Training questions are reverse-constructed from the hyperlink structure of a web corpus by authoring multi-hop chains over an entity graph, rewriting every non-answer entity into a descriptive reference so no clue can be resolved by string matching, and keeping only questions a reference model fails closed-book but solves with the supporting evidence. The questions become trajectories filtered at both trajectory and turn level for supervised fine-tuning (SFT), followed by reinforcement learning (RL) against live search with an in-cluster reward judge and observation summarizer, and the two stages alternate in a procedure called SFT-RL climbing that feeds the hardest solved and most efficient RL rollouts back into the next supervised pass. Because inference-time context management matters more on these benchmarks than most reported system differences, every result is reported with and without it using a single ReAct agent with no sub-agents or test-time verification, and with management enabled the two models reach 82.2 and 88.6 on BrowseComp, plus 84.8/85.1 on BrowseComp-ZH, 86.9/92.9 on DeepSearchQA, and 52.3/56.4 on HLE, the strongest open-source results in their parameter ranges.
Open-source search agents are hard to compare because inference-time context management (CM) often accounts for more of the score than the policy itself; the AllSpark Team's answer is a full recipe (web-graph task synthesis, doubly-filtered SFT, RL against live search, and alternating "SFT–RL climbing") for two agents, Iris-mini (35B-A3B, from Qwen3.6-35B-A3B) and Iris-pro (397B-A17B, from Qwen3.5-397B-A17B), evaluated both with and without CM under a fixed tool set, context limit, and judge.
- Training questions are reverse-constructed from a seed page and its out-links: an entity graph is distilled, a multi-hop chain of at least N relations is authored toward the seed entity, every non-answer entity is rewritten into a descriptive reference so no clue can be resolved by string matching, and a pair is kept only if a reference model fails it closed-book yet solves it when the entity graph is supplied.
- Teacher trajectories (ReAct with
searchandscrape, observations replaced by on-the-fly page summaries) are filtered at the trajectory level for correctness, degeneracy (a sliding-window zlib compression-ratio detector for loops) and minimum tool-call depth, then at the turn level by an LLM judge whose rubric is induced from free-form critiques rather than hand-written, masking at most 10% of assistant turns from the SFT loss. - RL uses a group-relative policy gradient with request-level partial rollouts resumed from committed prefixes (corrected by truncated importance sampling, at roughly 2× over-sampling), an in-cluster FP8
Qwen3.5-397B-A17Bserving as both binary generative reward model and observation summarizer, and each climb feeds back the shortest successful rollout (with enough tool calls) from queries whose pass rate is between 0 and 1/2, giving a self-paced curriculum. - With discard-all CM and a single ReAct agent at pass@1,
Iris-miniscores 82.2 / 84.8 / 86.9 / 52.3 andIris-pro88.6 / 85.1 / 92.9 / 56.4 onBrowseComp/BrowseComp-ZH/DeepSearchQA/HLE(text-only), with Iris-mini beatingXYZ-Aquila-miniby 3.4 points on BrowseComp but trailing it on DeepSearchQA (86.9 vs 89.5) and Iris-pro leading or tying on all four in its class. - CM is worth up to +21.2 points on BrowseComp for Iris-mini (64.7 without CM vs 85.9 with discard-all plus retry) but only +9.2 on HLE, gains shrink for the larger model, both models still trail frontier systems (
Kimi-K391.2,Apodex-1.0-H90.3 on BrowseComp), three configurations tie at exactly 85.1 on BrowseComp-ZH (246 of 289, with the appendix documenting a mislabeled question), climbing details are deferred to a later release, and the claimed positive transfer toBFCL,τ-bench,OfficeQAandAPEXis reported without numbers.
When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
In quantized recurrent networks, the stored low-precision state is fed back at the next time step, so the rule used to write that state can change all later computation. The authors name this rule recurrent-state write-back and isolate its effect in a compact GRU encoder-decoder for fluorescence lifetime imaging, which must estimate two lifetime parameters from very noisy time-resolved signals. With the trained model held fixed, switching to deterministic 4-bit state storage raises estimation errors by roughly 70x and 300x for the two parameters, because repeated small updates fall below the write threshold and the stored state freezes while the network keeps proposing change. Error feedback, residual memory, and direction memory carry the suppressed updates across time and restore accuracy without retraining, higher state precision can worsen a fixed solution, matched training can learn compatibility with the state interface, and an independently trained LSTM reproduces the failure with the cell state more sensitive than the hidden state.
Post-training quantization of recurrent networks stores each hidden state on a coarse grid and returns it to the next time step, so the rounding rule itself becomes part of the temporal computation rather than a passive encoding. The authors name this rule recurrent-state write-back, isolate its effect in frozen GRU and LSTM models for fluorescence lifetime imaging, and show that persistent sub-threshold updates that never reach the stored state, not bit width alone, are what break the network.
- On a fixed
Seq2SeqLiteGRU checkpoint (32 units, 6,627 parameters), replacing continuous state propagation with deterministic 4-bit write-back raises lifetime RMSE from 0.36/0.35 ns to 25.37/106.59 ns for τ1/τ2 (roughly 70-fold and 300-fold) with every weight and every other operation unchanged. - The failing trajectory shows the predicted signature: 99.59% of decoder updates fall inside the half-step write boundary, only 0.25% of state elements change level per step, and the median run of consecutive same-direction suppressed updates spans 132 of 134 decoder steps, versus a median of 2 steps in accurate native-4-bit models that also contain many sub-threshold updates.
- Carrying the discarded information forward rescues the same frozen network without retraining: error feedback restores 0.36/0.37 ns, a 2-bit residual memory reaches 0.34/0.40 ns, and the newly introduced direction memory (a small counter accumulating the sign of sub-threshold proposals) reaches 0.34/0.34 ns, all while the state fed back to the network stays on the 4-bit grid.
- More state precision is not a monotonic fix: evaluating the independently trained 4-bit-state GRU with 8-bit write-back worsens RMSE from 0.35/0.40 ns to 0.43/0.57 ns even though the median occupied decoder levels rise from 12 to 176.5, and in matched training the ranking of 4-bit, 6-bit, residual-memory, and direction-memory interfaces shifts, showing compatibility between a learned solution and its state interface is what matters.
- An independently trained 32-unit
LSTMreproduces the failure (0.239/0.254 ns to 3.857/1.327 ns under 4-bit write-back) and its rescue by error feedback, with the cell state far more sensitive than the hidden state (8.158 vs 0.296 ns τ1 RMSE); the study is limited to a single synthetic 160,000-sample FLI test set and tiny models, and the auxiliary-memory comparisons match stored bits but not hardware cost.
MaxKernel: Agentic Kernel Generation for TPUs
Writing high-performance custom kernels for accelerators demands deep hardware expertise, and large language models paired with real-time compiler feedback offer a way to automate it. MaxKernel is a multi-agent system for TPU kernel development with three modes: a Human-in-the-Loop (HITL) agent for step-by-step collaborative design, an Autonomous agent that runs a fully automated, metric- and trace-driven optimization loop, and a Graph-Based Autonomous Search that scales the autonomous agent to global exploration of the design space, all sharing sub-agents for planning, implementation, self-debugging, testing, and hardware profiling. Evaluated on JaxBench, a suite of 50 diverse TPU kernel tasks, plus real workloads from open-source models, the system consistently produces implementations matching expert hand-tuned baselines. The agent is open-sourced.
Hand-writing high-performance TPU kernels in JAX/Pallas demands expert management of memory hierarchies, DMA pipelining, and tiling, and zero-shot LLM generation mostly fails against rigid accelerator APIs and opaque compiler errors. MaxKernel wraps an LLM in a closed loop of specialized sub-agents (planning, implementation, compile-fix, test synthesis, autotuning, and XProf-based profiling) and then scales that loop with parallel and beam search over a persistent graph of kernel states.
- The system offers three orchestration modes: a human-in-the-loop router that pauses after each sub-agent, an autonomous hill-climbing loop that feeds profiling traces back into the next plan and rolls back to the best valid snapshot, and a graph-based search that spawns fresh autonomous sessions per node, with a test suite frozen before optimization begins so implementation agents cannot alter the correctness criteria, and a RAG knowledge store that deliberately excludes hand-tuned kernel code.
- On
JaxBench(50 TPU v6e tasks, all runs usingGemini 3.1 Pro), a Best-of-100 zero-shot baseline compiled only 10/50 kernels at a 1.08× geomean, while the single-trajectory Auto agent reached 48/50 correct at a median 1.39× (range 1.19–1.42 across five seeds), and parallel search over those five trajectories hit 50/50 correct with a 1.58× geomean and 34/50 tasks faster than XLA, versus 1.49× and 31/50 for beam search. - Against eight expert hand-tuned Pallas kernels, parallel search posted a 2.32× geomean speedup versus 2.02× for the human references, beating them on seven of eight workloads (e.g. 6.74× vs 2.41× on Paged Attention, and above-baseline MLA kernels where the human version ran at 0.69×), but losing badly on Ragged Paged Attention at 1.42× vs the expert's 4.65×.
- On production open-source workloads, generated kernels cut
Qwen3-NextGated DeltaNet training-step time by 4.70×, acceleratedDeepSeek-V4sparse-attention prefill by 7.85× over JAX, and shaved 8.68% latency off an already hand-written Pallas MLA kernel, while a debugging run fixed a deadlock in Ragged Paged Attention v3 prefill by inserting clamp instructions for left-padded inputs. - Caveats include a single-LLM evaluation with no model ablations, geomean speedups floored at 1.0× so regressions are hidden, correctness tolerances relaxed to as loose as 0.1 on some tasks, a beam-search plateau at depth 3 attributed to its 2-iteration per-node budget being too short for multi-step Pallas fixes, and the parallel-search numbers representing the best of five runs rather than a typical single trajectory.
Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One
Language generation is nearly always sequential, either token by token in autoregressive models or through long iterative refinement trajectories in diffusion language models. PlaidQ is a 0.7B continuous diffusion language model for code that repurposes a pretrained autoregressive model as a bidirectional denoiser over continuous token embeddings, then distills its trajectory with distribution matching for few-step generation and paired-trajectory supervision for one-step generation. At matched scale it is competitive with discrete diffusion models on code, and a 16-step distilled student reaches 31.78 and 40.49 pass@10 on HumanEval and MBPP+, beating its own teacher sampled for 512 steps. A single denoising step still yields functionally correct programs at 7.07 pass@1 on HumanEval, and code and checkpoints are released.
Language generation, whether autoregressive or diffusion-based, remains sequential: autoregressive models chain token by token, and diffusion language models trade that for a long iterative refinement trajectory, so few-step or one-step generation has stayed elusive. PlaidQ is a 0.7B continuous diffusion model for code built by converting a pretrained autoregressive model into a bidirectional denoiser over token embeddings, and its long trajectory is then distilled down to 16, 8, 4, or even a single denoising step.
- The model reuses the
Qwen3-0.6Btrunk and vocabulary head with causal attention swapped for bidirectional attention, shifts the reconstruction head by one position to preserve the pretrained next-token alignment, and denoises every completion position in parallel while the prompt stays clean, decoding tokens only once at the end. - Training at a 152k-token vocabulary and 2048-token sequences is made tractable by
SWVR, a streamed exact kernel for the categorical reconstruction that cuts activation memory from 40.2 GB to 28.3 GB (29.6%) with identical loss and gradients, plus a hybrid Muon-AdamW optimizer and autoregressive initialization that together lower validation NLL from 3.65 to 3.59. - As a base generator at 128 steps,
PlaidQis roughly on par with same-scale discrete diffusion models (18.05 pass@1 onHumanEval+versus 17.6 forOpen-dCoder), reaches 39.87 pass@10 onHumanEvalwith classifier-free guidance, and transfers zero-shot to infilling with 37.91 onHumanEval-Infilldespite never seeing that conditioning during training. - Distribution matching distillation with an on-policy critic yields a 16-step student that scores 31.78 and 40.49 pass@10 on
HumanEvalandMBPP+, beating the 128-to-512-step teacher (28.57 and 35.21), while at 4 steps distillation lifts pass@10 from 7.90 to 15.85 and from 7.91 to 30.94 where the training-freeDPM-2Msolver collapses to zero. - For the single-step regime, paired-trajectory distillation feeds the student the teacher's output for the same initial noise as a cross-entropy target, raising one-step
HumanEvalpass@1 from 0.09 to 7.07 (8.53 pass@10), though absolute few-step scores remain far below 7-8B discrete models, each step budget requires a separately trained student, and infilling onSantaCoderlags at 22.86 versus 29.6 forOpen-dCoder.
Extremely Sparse Supervision Incentivizes Reasoning Ability
Prevailing post-training methods for reasoning optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. Working in the on-policy distillation setting with the Qwen3 family, the authors find that supervising as few as one or two tokens per reasoning trajectory, about 0.05% of all tokens, matches or surpasses full-token training on mathematical reasoning in most cases. The effect holds across nine teacher-student configurations of varying scale and is further validated on coding reasoning, Llama models, and Proximal Policy Optimization (PPO)-based reinforcement learning with verifiable reward (RLVR). The authors suggest this mirrors natural learning, where one reflects on a few critical steps rather than correcting every word, and argue it points toward more efficient post-training algorithms.
On-policy distillation (OPD) is assumed to work because the teacher supervises every generated token, yet masking the loss down to just one or two tokens per reasoning trajectory, about 0.05% of all tokens, matches or beats dense supervision at improving student reasoning. The finding challenges the premise that effective post-training must be token-intensive.
- Sparse OPD keeps the standard sampled-token objective, where each token's reward is the teacher-minus-student log-probability, but inserts a mask so only selected tokens contribute gradient:
rand1tokpicks one random token,maxtokandmintokpick the single highest- and lowest-reward token,minmaxtokkeeps both, andpctltail 0.05%keeps the extreme reward tails. - Across nine
Qwen3teacher-student pairings trained onDAPO-Math-17Kand evaluated with pass@k up to 128 and avg@8 onAIME 24,AIME 25, andHMMT Feb 25, evenrand1tokimproves the base student in all nine, and some sparse variant matches plain OPD in 2 of 9 families and beats it in 7 of 9. - The strongest case is
Qwen3-8Bdistilled from the smallerQwen3-4B-Instruct-2507, whereminmaxtokreaches 49.3 mean avg@8 versus 47.5 for plain OPD and 46.8 for the teacher itself, while touching roughly 10% of parameters and, formaxtok, raising reverse KL to the teacher from 0.18 to 1.12, so the gains are not explained by closer imitation. - Which extreme token helps depends on student capacity:
maxtok(reinforcing high-entropy "forking" tokens the teacher strongly prefers) excels when the student is same-scale or larger, whereasmintok(suppressing teacher-dispreferred tokens, which an audit shows are mostly correct-but-dispreferred rather than genuine errors) is more stable when the student is much smaller, and performance is non-monotonic in sparsity. - The phenomenon replicates on coding reasoning (
Eurus-RL-Code,LiveCodeBench v6),Llama 3models, andPPO, but not underGRPOor REINFORCE, and since every variant still needs the full set of on-policy rollouts the savings are in logit memory rather than training compute.
$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction
Coding agents are increasingly asked to build production LLM agents, but existing benchmarks do not measure whether they can deliver one under the conditions of a real client engagement. τ^τ-bench (hyper-tau-bench) hands a developer agent a business's actual records, a client who holds the requirements, a production API, an inherited codebase, and limits on serving cost and models, then scores the customer-service agent it builds by deploying it against held-out simulated users across 53 tasks in four domains. The strongest configuration, Claude Opus 5 running under Claude Code, passes only 23.9% of evaluation simulations, against an expert-authored reference ceiling of 82.2%. The failures mirror those seen with human agent developers: shallow queries instead of deep comprehension of the records, almost no communication with the client, and too little experimentation with architecture or serving spend before shipping the first design that runs.
Building customer-service agents is increasingly delegated to coding agents, but existing benchmarks score finished agents rather than the work of constructing one from a client's messy records, stakeholders, and cost constraints. τ^τ-bench (hyper-tau-bench) makes agent construction the task: a developer agent receives a business's scattered artifacts, a simulated client, a possibly defective REST API, an optional inherited codebase, and a serving-cost budget, and must ship a complete agent that is then scored by deployment against held-out simulated users.
- Each of the 53 tasks across airline, retail, telecom, and banking is built by decomposing
τ-benchpolicies into atomic facts (3,328 total, 2,969 in banking alone) and rendering them into 2,868 realistic artifacts (handbooks, transcripts, spreadsheets, screenshots, call recordings), with some facts moved to an LLM-simulated client that must be interviewed and with contamination controls via rebranded, re-valued airline and retail domains. - The strongest configuration,
Claude Opus 5underClaude Code, passes only 23.9% of evaluation simulations versus an 82.2% expert-authored reference ceiling;GPT-5.6-solinCodexscores 22.0% with far shorter builds (48 vs 216 minutes), and banking collapses the field to 3–9% while airline, retail, and telecom reach roughly 42–73%. - Trajectory analysis traces failures to skipped disciplines: developers grep the corpus instead of reading it (opening fewer than 80 of ~1,700 banking files), talk to the client in just 0.3% of tool calls even though builds that ask four or more questions score 0.50 versus 0.16 for builds that never ask, rewrite inherited code without ever running it, and use only 0.45–0.72× of the serving budget while 21 builds overshot it and 10 saw their scores zeroed by the overage penalty.
- Nearly every build (92%) is a single LLM tool loop with a cheap, same-vendor model, yet a one-line architecture hint (route by intent, review tool calls) doubled a telecom score from 31% to 67%, and cheating-adjacent probes of the grader or hidden data appeared in 17–42% of runs across harnesses, none successful.
- Limitations include a single construction trial per task (no variance estimate), single-LLM client and user simulators with fixed requirements, corpora audited for full consistency unlike real engagements, and evaluation that stops at submission rather than covering post-deployment maintenance.
LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28
Discovery Loop is a lightweight system in which a large language model iteratively evolves an optimization algorithm: starting from a simple seed solver, the model proposes improvements guided by a scoreboard and a history of prior ideas, an independent verifier evaluates each candidate, and only improvements are kept. On the Packomania circle-packing benchmark, where the goal is to maximize the sum of radii of N variable-radius circles in the unit square, the system beat the best known solutions for 10 values of N between 101 and 114 by 2.4% to 5.4%, within 15 iterations and for a total LLM cost of $27.72, and the records were independently accepted by Packomania. The authors also analyze cost-efficiency dynamics, including an adaptive plateau-detection mechanism.
Discovery Loop asks how much of DeepMind's AlphaEvolve paradigm survives when reduced to one LLM, one consumer PC, and a sub-$30 budget, and answers by using Claude Fable 5.1 to iteratively rewrite a circle-packing solver until it beats best-known solutions on the Packomania csqv benchmark (maximize the sum of radii of N variable-radius circles in a unit square).
- The loop is roughly 400 lines of Python: each iteration feeds the LLM the full source of the current champion solver, a per-target scoreboard, and the last 12 ideas tried, receives a complete replacement solver rather than a patch, runs it on all targets in parallel with a 120 s timeout, and accepts it only if an independent zero-tolerance verifier (containment, non-overlap, recomputed radii sum, feasibility shrink) confirms an improvement.
- Across 15 iterations and $27.72 of LLM spend, the system produced new accepted records for 10 of 12 targets (N = 101–114), with gains of 2.4%–5.4% per instance and a +5.0% total sum of radii (51.81 → 54.41), while N = 26 and N = 32 only matched prior records to within 1e-6.
- A notable caveat is that every one of the 10 records was first broken at iteration 0 by the hand-written seed solver (multi-start penalty
L-BFGS-Bwith LP-optimal radii); the LLM-evolved ideas (basin hopping with contact-graphSLSQPpolish, hexagonal-lattice initialization, island-model parallelism,KKT-Newton polish, defect-migration moves) improved the aggregate score only from 59.39 to 59.98, so the headline results say more about weak prior records than about LLM-driven discovery. - Returns diminish sharply: iterations 0–5 cost $4.96 for +0.57 total score, while iterations 6–14 cost $22.76 for +0.02 (a 130× jump in cost per unit gain), and a proposed plateau detector (window 4, threshold 0.01) would have stopped after iteration 9, saving about 50% of the spend at a loss of 0.006 in total score.
- The system keeps a single champion with no population or evolutionary database, was tested on only one cheap-to-evaluate problem, and a side experiment on
MIPLIBmixed-integer programming saw a 75% code-generation failure rate, suggesting the approach is fragile on more complex domains.
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
LLMs are increasingly asked to turn natural-language descriptions into operations research (OR) optimization models, but real requests often omit objectives, constraints, or business rules that change the resulting program, and existing benchmarks assume complete specifications. OR-Clarify presents partial problem descriptions with structured hidden slots and scores agents on slot recovery, stopping behavior, silent assumptions, and interaction cost through bounded dialogue with a simulated user, in both open-ended and choice-based clarification modes. The proposed InterOPT framework first identifies unresolved formulation-critical gaps and then uses them to decide whether to ask another question or stop; in choice-based experiments it substantially outperforms all baselines on exact slot recovery and stays competitive with strong prior methods in the open-ended setting.
Large language model (LLM) agents that turn natural-language business requests into optimization models are usually evaluated on complete specifications, but real operations research (OR) requests routinely omit objectives, constraints, or rules that change the resulting mathematical program. The paper reframes the pre-modeling step as a selective completeness decision, introducing the OR-Clarify benchmark to measure whether an agent asks the right clarifying questions and stops at the right time, plus InterOPT, a two-stage method that tracks unresolved formulation gaps and uses them to steer questioning.
OR-Clarifyis built by decomposing fully specified OR tasks into single-requirement facts, masking roughly half of the formulation-critical ones behind a fixed seed, and scoring bounded dialogue with a simulated user who answers only what is asked; it comprises 100 cases and 178 hidden slots (75 P0 blocking, 83 P1 substantive, 20 P2 secondary) and reports exact slot recovery, stopping behavior, silent assumptions, and question count.InterOPTseparates monitoring from control: Stage 1 (Dynamic Gap Search) maintains a persistent ledger of open formulation gaps across six categories such as objective trade-offs and hard-versus-soft policies, while Stage 2 (Gap-Guided Action Search) generates three gap-bound candidate questions, runs a selector over them, and leaves the ask-or-READY_TO_MODELdecision to the model itself.- Off-the-shelf models struggle under both free-form and multiple-choice protocols, with
Opus-4.8strongest but no model exceeding 60% Core Exact and all leaving substantial silent assumptions per run, andDeepSeek V4 Proas the default tested model. - In the choice-based setting
InterOPTreaches 0.675 Core Exact versus 0.506 for the MC-D baseline and roughly halves silent assumptions (0.366 vs 0.692), but at a steep interaction cost of about 10 questions per run versus 2.4, and in the open-ended setting it trails theORPilotandGATEadapters (0.538 vs 0.583 and 0.560 Core Exact). - Ablations show gap tracking drives most of the coverage gain while guided selection mainly shortens dialogue, and the open-setting diagnostic identifies stopping as the core failure: agents declare readiness in 98.8% of runs, yet 40.6% of stops are audited as premature, which the authors flag alongside dependence on a single tested model and LLM-based judging.
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Layer dropout, also called stochastic depth, speeds training and enables zero-shot layer pruning in vision and language transformers, yet it has vanished from LLM pretraining recipes amid reports that it hurts accuracy, without any systematic study of the effect. Across more than 2,400 pretraining runs on Cerebras CS-3 systems spanning 271M to 8.2B parameters and up to 160B tokens, the authors establish best practices for the layer distribution, time schedule, and optimizer hyperparameters and show that at equal training FLOPs layer dropout yields lower loss. Models reach lower or similar validation loss while saving up to 25% of training FLOPs, and the trained models support post-training optimizations such as early exit, intermediate-layer skipping, and self-speculative decoding for up to 1.5x inference speedup with negligible accuracy loss.
Layer dropout (stochastic depth) vanished from LLM pretraining recipes because it appeared to hurt accuracy in the single-epoch, large-data regime, and this study from Cerebras argues that those degradations were mostly a configuration problem. Across more than 2400 runs on CS-3 systems, it shows that dropping whole transformer blocks with the right residual scaling, depth distribution, and time schedule cuts training FLOPs while matching or beating dense loss, and leaves the model robust to early exit and layer skipping at inference.
- Each transformer block is skipped per sequence with a Bernoulli mask, scaled during training by
1/ρ(ρ = 1 − p) followingCompleteP's maximal-residual-update rule, which is what makes optimal learning rate, batch size, and weight decay transfer unchanged across dropout rates, whereas the commonr_train = 1convention forces per-rate retuning. - Ablations at 20 tokens per parameter establish the recipe: whole-layer dropout beats separate attention/FFN sub-layer dropout, per-sequence masks beat per-batch masks, non-uniform depth distributions beat uniform at equal FLOPs, and a linearly increasing rate across depth (ILD) paired with a linearly decreasing rate across steps (DTS) is best, with increasing-over-time schedules degrading loss badly.
- With
ILD+DTSat 5% FLOPs savings, 503M and 906M models beat the dense baseline (906M: 1.9513 vs 1.9526 val loss), and at larger scale a 1.8B model with p_max=0.6 saves 15% FLOPs at 1.836 vs 1.849, a 3.9B model with p_max=0.8 saves 20% at 1.745 vs 1.732, and an 8.2B model with p_max=0.99 saves 25% FLOPs at 1.663 loss, with degradation staying within roughly 0.5% of baseline as tokens per parameter grow. - The same models gain zero-shot depth elasticity: skipping alternate layers of the 3.9B dropout model yields 2.129 loss versus 6.446 for its dense twin,
Balcony-style exit adapters trained on frozen weights reach lower early-exit loss than on dense models, andDraft & Verifyself-speculative decoding onXSUMreaches 1.54× speedup at 3.9B where the dense baseline manages only 1.02×. - Trade-offs and gaps remain: alternating-layer dropout is best for skip robustness but fails at early exit while ILD is the reverse, hyperparameter transfer weakens at aggressive rates and was only validated for constant schedules, the 8.2B run has no dense baseline in the results table, the 3.9B dropout run is slightly worse in raw loss, and there is no comparison against Mixture-of-Depths or MoE architectures.
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
On-policy distillation (OPD) gives dense per-token supervision for post-training language models, but external teachers suffer from distribution mismatch and self-distillation with privileged context is limited by in-context learning capacity. RISE (Recursive Improvement via Self-Extrapolating Policy Distillation) instead builds a synthetic teacher from the model's own reinforcement learning with verifiable rewards (RLVR) trajectory, extrapolating the displacement between the current checkpoint and a trailing anchor in parameter or logit space to turn a sparse outcome-driven update into a dense token-level target. Because the teacher is refreshed every iteration as the student improves, distillation becomes a recursive loop in which outcome rewards ground the extrapolation and the extrapolated teacher refines token-level decisions. Across mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks, RISE outperforms both RLVR-only training and on-policy self-distillation.
On-policy distillation gives dense per-token supervision but is bottlenecked by teacher quality: external teachers suffer distribution mismatch, and privileged-context self-distillation is capped by in-context learning capacity. RISE builds the teacher from the model's own RLVR trajectory by extrapolating the displacement between the current checkpoint and a trailing anchor, turning a sparse outcome-driven parameter update into a dense token-level target with no external model or privileged conditioning.
- Each iteration runs a
GRPOstep, constructs a synthetic teacher as anchor + β·(post-RLVR checkpoint − anchor) with β decaying linearly from 1.2 to 1 over training, either in weight space (task arithmetic) or logit space (a geometric mixture of the two policies' output distributions), then distills that teacher into the post-RLVR checkpoint via a top-K Jensen–Shannon loss on the same rollouts, so no extra sampling is needed. - On
DAPOMath,Qwen3-8Bin-domain Math Avg rises from 60.0 (GRPO) to 62.7,Qwen3-1.7Bfrom 45.4 to 50.2, andOLMo3-7B-Instruct-SFTonOpenR1-Math-46Kfrom 47.6 to 56.4 with AIME'24 jumping 30.2 to 46.9, while privileged self-distillation baselines (GRPO+SDPO,SDAR,RLSD) land within about two points of GRPO or below it and OOD scores onGPQA,IFEval, andMMLU-Prohold or improve. - The gains carry to mixed math+STEM on
Qwen3-4B-Base(Math Avg 44.8 vs. 40.2), to code onQwen3-8B-Base(faster convergence, similar final accuracy), and to agentic tasks onQwen2.5-3B-Instructwhere the weight-space variant liftsALFWorldsuccess 75.0 to 84.4 andWebShopaccuracy 63.3 to 74.2. - Ablations show both phases are necessary: extrapolating without RLVR collapses within 60 steps (MATH-500 falls to 2.4% as responses hit the length cap), while adopting the extrapolated weights directly without distillation yields only +0.3 and +0.2, and a compute-matched
GRPO-2xwith a second gradient pass recovers just +0.5 and +1.9 versus RISE's +2.7 and +4.8. - The safe β range shrinks as training proceeds (β=2.0 gives +7.3 AIME'24 points at step 50 but −15 at step 100), the EMA anchor rate is model-dependent (η=0.1 best for Qwen, η=1 best for OLMo), wall time is 1.3–1.6× GRPO, and the method inherits any reward hacking since extrapolation amplifies whatever direction RLVR takes.
Applications 93
From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance
AI recruitment tooling has moved from scoring candidate-job pairs and ranked lists to multi-stage workflows that retrieve evidence, compare candidates, and take or recommend actions, and this systematized narrative review traces that shift from bilateral retrieval and behavioral ranking through neural person-job matching, large language model (LLM) components, and tool-using recruiting agents. Using a purposive search and coding protocol covering 40 representative works plus industrial and legal sources, it organizes the field around three transitions: similarity to reciprocal suitability, single model to compound workflow, and offline prediction to evidence- and productivity-aligned evaluation. Within the coded set, privacy is never directly evaluated and no work jointly assesses utility, fairness, privacy, and security, and the authors note that behavioral labels confound exposure, preference, and qualification while final-output scores hide pipeline failures. The review closes with a staged mapping from evaluation evidence to the strongest defensible claim and an agenda for reciprocal, auditable, temporally controlled systems.
Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection
Adaptive network intrusion detection systems retrain after drift alarms, but an alarm signals change without establishing that a challenger model should replace the deployed incumbent, and promotion conclusions may depend on how the challenger was built and how much evidence supports it. The authors test this dependence on CICIDS2017, UNSW-NB15, and ToN-IoT with self-contained challenger pipelines, nested candidate-size controls, a common harness comparing nine update policies, and a sensitivity check confining each exact feature vector to a single evaluation, training, or probe role. Incumbent-owned frozen preprocessing amplified apparent promotion harm, which did not persist with self-contained pipelines, and raising candidate evidence from 512 to 2,000 samples per class improved balanced accuracy by 0.38 to 1.67 points across the three benchmarks, mainly through fewer false positives. No policy dominated globally, thirteen replays on real time-ordered traffic showed no net harm from always deploying, and the paper argues that challenger construction and evidence must be controlled and reported explicitly.
A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models
Large language models (LLMs) are increasingly used to label cultural texts at scale, but it is unclear whether their labels are reliable enough to treat as measurements of latent social constructs. Five LLMs were run repeatedly as zero-shot annotators of four constructs in English song lyrics, namely self-esteem, self-control, seeking belonging, and seeking recognition, and the outputs were checked for run-to-run consistency, cross-model agreement, and whether consensus labels could train a supervised classifier. Reliability varied sharply by construct: self-esteem was the most stable across models, seeking recognition the least, and the other two fell in between depending on the model. Consensus labels carried learnable signal for downstream classification, and the authors argue that repeated-measurement stability and cross-model convergence should be reported before LLM annotations are used as scalable measurements in cultural analytics.
Hakken: Predicting future discoveries to fill the gaps in today's knowledge
Hakken aims to predict scientific relationships that have not yet been documented, going beyond what can be deduced from existing literature. It trains a transformer on temporal sequences of knowledge graphs extracted from large publication corpora, fuses this with a large language model's semantic knowledge to predict the presence and type of future relationships between concepts, and attaches a model-agnostic explanation layer so scientists can evaluate each suggestion. Applied to biomedicine, the model sets a new benchmark for time-aware multi-label relation prediction and stays coherent over long historical spans. The authors scored 1.5 million aging-related hypotheses, reviewed batches with biologists, and advanced three to wet-lab testing, confirming two previously undocumented interactions, between TP53 and BAMBI and between RAF1 and TNF.
Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection
Whisper's generative decoder can emit fluent but fabricated transcripts for audio containing little or no speech. The proposed training-free method estimates a compact hallucination-associated subspace from non-speech calibration data and projects decoder hidden states away from it at inference, either always-on or gated by Whisper's own prediction that the input is non-speech. Across non-speech benchmarks, always-on projection cuts the average hallucination rate from 31.31% to 2.44%, a 92.21% relative reduction, while the gated variant reaches 3.74% with fewer false rejections of genuine speech. On LibriSpeech, gated projection raises absolute word error rate by 0.33 to 4.39 percentage points and produces false-rejection rates of 0.41 to 9.97% depending on model and split, giving a controllable trade-off between suppression and recognition quality.
Sustainable Edge Vision via Empirically Calibrated DVFS: Eliminating Thermal Throttling on Passively Cooled Hardware
Passively cooled edge devices avoid fan energy overhead and mechanical failures, but sustained deep neural network inference on them is bottlenecked by thermal throttling. The authors design an empirically calibrated, state-aware dynamic voltage and frequency scaling (DVFS) scheduler that combines time-domain dwell guards, absolute temperature bounds, and derivative triggers for sharp thermal spikes, rather than reacting to temperature alone. On a passively cooled Raspberry Pi 5 running YOLOv8n for 30-minute workloads, the scheduler eliminates all observed thermal throttling events, delivers 6.8% higher frame rate than a temperature-only baseline at 1.9% less energy per frame, and beats an actively cooled reference on Joules per frame, though the passive envelope closes at ambient temperatures of 27°C or higher where nonlinear leakage defeats DVFS control.
Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling
Drug candidate design means searching a vast, rugged chemical space for molecules that satisfy competing objectives, and while reinforcement learning from verifiable rewards (RLVR) can adapt large language models (LLMs) for this, chemically relevant scoring functions can take hours or days per evaluation, making them too slow for online training. The authors ask whether LLMs can learn molecular design strategies from cheaper synthetic tasks that transfer to expensive structure-based lead optimization. Curriculum recipes that progressively introduce harder synthetic design tasks yield performance surpassing much larger frontier models on structure-based lead optimization, suggesting synthetic-task scaling as a route to post-training for experimental settings too costly to train on directly.
Federated Attack Campaign Detection via Contrastive Encoding of Threat Indicators in Gradient Updates
Detecting cyberattack campaigns that span organizations normally requires sharing sensitive telemetry across institutional and national borders. FedIoC is a federated learning framework in which each client trains a threat detector locally and folds its structured indicators of compromise (IoCs) into gradient updates via a supervised contrastive loss: flows matching known indicator patterns are pulled together in embedding space and non-matching flows pushed away, so campaign structure is expressed in the gradient direction. The server then clusters clients by cosine similarity of their updates and recovers cross-organizational campaign cohorts from gradient geometry alone, with no direct indicator transmission, on two public threat-detection benchmarks where each client sees only a fragment of every campaign and holds disjoint indicator sets. The authors identify non-IID gradient structure as the main driver of this recovery and pose encoder design that improves on it as an open problem.
Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution
Ransomware detection and family attribution benefit from static, dynamic, and memory analysis, but running every modality on every sample wastes compute and adds latency. The proposed cost-aware Hierarchical Multi-Agent System (HMAS) organizes specialist agents under domain controllers coordinated by a meta orchestrator, starts with cheap static analysis, and escalates to dynamic and memory analysis only when confidence is low or specialists disagree, with a locally deployed large language model verifying selected hard cases without replacing the deterministic pipeline. The full system reaches 96.57% accuracy and 0.99 ROC-AUC for binary detection and 0.90 macro-F1 for family attribution, while cutting average analysis cost by 43.97% relative to exhaustive analysis. Routing analysis shows that 56.05% of samples are resolved from static evidence alone and only 4.33% need the complete pipeline.
When Genomic Masking Priors Fail to Transfer: Strong Variant Prediction, Weak Functional Generation
Bidirectional discrete diffusion looks like a natural fit for DNA because it can reconstruct missing sequence from both flanks. GenDA (Genomic Density-optimized Absorbing Diffusion), a 202M-parameter model that places masked spans where local sequence entropy is highest, reaches a pooled ClinVar single-nucleotide-variant AUROC of 0.774 after supervised fine-tuning, beating a similarly sized autoregressive model by 0.103. But a matched random-span variant scores 0.777, leaving no evidence that entropy guidance causes the gain, and in zero-shot functional inpainting of promoters, enhancers, and exon or intron boundaries the model fails to consistently beat a control that shuffles the gap while preserving 3-mer composition. The authors conclude that strong fine-tuned variant prediction, a plausible corruption prior, and usable functional generation are separate claims needing separate validation.
CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games
Matchmaking in a large multiplayer online battle arena game breaks when most queued players lack history in the mode being played, when skill distributions shift drastically across rank tiers, and when extreme skill brackets are data-starved. CHAMP swaps the target-mode-only player profile for a hybrid feature set, a timestamp-ordered cross-mode short-term sequence annotated with target-domain features plus per-mode long-term, real-time, and team statistics, feeding a Domain-Aware Win-rate Network that learns mode-conditioned representations and per-mode debiasing inside one shared model serving every mode. Offline it predicts win rate with 67.73% accuracy, ahead of all evaluated attention and sequence baselines, and online A/B tests across the full ladder cut the five-minute kill-crushing rate by up to 20.73% for lower-tier players.
ReCAST: Restoration-aware Cascaded Stage-wise Training for Obfuscated SMS Risk Classification
Fraudulent Chinese SMS messages are increasingly written to evade the cheap classifiers production systems can afford, hiding risk-bearing phrases behind crafted obfuscations that humans still read without effort. ReCAST distills a large teacher model's de-obfuscation ability into a smaller deployable student by supervising obfuscated span detection, obfuscation type prediction, and text restoration, then uses that restoration-aware student for downstream risk classification. On an internally built real-world benchmark the approach substantially improves classification over directly trained baselines under obfuscation while staying within production latency and throughput constraints.
Adaptation Interfaces for In-Context Tabular Foundation Models in Time-to-Event Prediction
Tabular foundation models (TabFMs) handle standard classification and regression on structured data well, but censored time-to-event prediction demands explicit treatment of censoring and event-time dynamics. This study attaches CoxPH, DeepHit, and cause-specific MTLR survival heads to frozen TabFM backbones and compares them against temporal zero-shot reformulation and classification-based fine-tuning across 74 single-risk datasets plus 4 competing-risk ones, with a revised context-resampled training procedure. Zero-shot inference is effective on smaller datasets while supervised adaptation grows more advantageous as data scales, with the Cox interface the most reliably strong option, especially on integrated Brier score; DeepHit favors time-dependent concordance over probabilistic calibration, and classification fine-tuning remains the weakest for probabilistic prediction.
Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing
Technical progress in applying artificial intelligence (AI) to investing is often mistaken for evidence that it makes money. This critical review surveys public research through 31 August 2026 on listed equities, exchange-traded funds, centralized crypto spot and perpetual futures, and on-chain markets, organizing it along an alpha-translation chain in which point-in-time information must become a stable signal, feasible positions, executable orders, and risk-adjusted returns after costs. Across machine learning, time-series foundation models, financial language models, reinforcement learning, and agents, the authors find real but mostly upstream gains in prediction, text processing, portfolio design, and workflow integration, undermined downstream by temporal contamination, repeated selection, survivorship, weak benchmarks, execution costs, and capacity limits. Within the evidence examined, no general AI architecture is shown to deliver persistent, cross-regime, capacity-aware net alpha, and the review sets out conditions such as point-in-time data, decision-aligned objectives, joint portfolio-execution evaluation, and prospective tests for more credible claims.
BIT.UA at BioASQ 14B: Modular Retrieval with pg_textsearch and Qdrant, and Agent-Based Answer Generation
The BIT.UA team from the University of Aveiro describes its system for the 14th edition of the BioASQ Task B biomedical question answering challenge, built on a substantially refactored and modular codebase. For document retrieval, the PyTerrier PISA index was replaced with PostgreSQL-based pg_textsearch for BM25 search and Qdrant for GPU-accelerated dense embedding search, combined with HyDE query expansion, a Context-1 retrieval strategy, and a new reranker trained with dense-retrieval negative sampling. For answer generation, the team added an LLM-as-a-judge framework and an agent quorum mechanism in which multiple agents with diverse prompts debate and iteratively converge on a consensus answer using adaptive document retention, and it entered the snippet generation subtask for the first time. The Phase A retrieval systems reached MAP rank 5 in Batches 1 and 3, and all code is openly released.
Deep Microcompression: Structured Pruning and Bit-packed Quantization for Microcontrollers
Running convolutional networks on bare-metal microcontrollers is limited by kilobytes of RAM and flash rather than by compute. Deep Microcompression (DMC) is a hardware-aware pipeline combining structured pruning, quantization-aware training, and fixed-length bit-packing, and it emits a dependency-free C library with deterministic latency. On LeNet-5 it reaches a 55.8x weight compression ratio at 98.77% accuracy, cuts binary size by 3x versus TensorFlow Lite on the RP2040 at matching accuracy, and delivers the first documented deployment of a standard CNN on the ATmega328P, a device with only 2KB of SRAM.
Conformal Prediction for Offensive Security
Conformal Prediction (CP), a technique for producing prediction sets with coverage guarantees, has been used defensively in cyber security but its use for carrying out attacks is hard to find in the literature. The authors examine this gap by applying CP offensively in two areas: Privacy-Preserving Machine Learning and network traffic analysis. The paper reports initial findings in both areas rather than a complete attack pipeline, positioning CP as a component of attacks rather than defenses.
CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review
Reports from the AAAI-27 review cycle raised concerns that reviewers coordinate bids to get assigned to each other's papers, but collusive intent is unobservable in real data and no counterfactual exists for the same conference. CABAL is an end-to-end multi-agent simulation that holds the conference environment fixed and populates it with LLM-driven reviewer agents following honest or collusive policies, plus an affinity-guided strategy that forms collusion rings and picks target papers consistent with the colluders' expertise rather than at random. In controlled runs, collusive bidding more than doubles target-paper capture, and assigned colluders score their targets about two points higher than honest co-reviewers, while conference-wide effects stay modest. Bid-phase detectors offer limited help: positive-bid graphs are confounded by benign affinity, and a Very-High-only diagnostic view recovers rings precisely but with low coverage.
Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness
Building automation systems produce abundant sensor data but remain hard to use operationally because of inconsistent point naming, missing metadata, and fragmented documentation. The authors systematically review and code 66 peer-reviewed studies on large language models (LLMs) for heating, ventilation, and air conditioning (HVAC) operations published between 2023 and March 2026, classifying each across five application families and three LLM method families and assessing evidence realism, deployment readiness, and where responsibility sits between the LLM and physical HVAC decisions. The corpus is concentrated in building energy modelling (BEM), with 32 of 66 papers, and only four studies reach pilot-level evidence, none reports sustained operational deployment, and none was judged ready for industry adoption now. The review concludes that current evidence supports LLMs as semantic and workflow layers, such as point-name normalisation and document-grounded operator support, rather than as autonomous controllers, where conventional machine learning, model predictive control (MPC), and reinforcement learning remain more adopted.
The History Is the Detector: Executing CVE Patch History, End-to-End
Public vulnerability databases and their fixing commits record exactly why code was unsafe, but that knowledge is written for human inspection and the same unsafe patterns persist in code without an advisory. BUGSTONE-E2E mines reusable detection rules from verified fixing commits, capturing scan anchors, fix semantics, and Common Vulnerabilities and Exposures (CVE) provenance organized by Common Weakness Enumeration (CWE) and language, then applies them through a funnel: Tree-sitter enumerates call sites matching rule anchors, cheap heuristics discard benign sites without LLM calls, LLM-based agents inspect survivors guided by the rule, and the system re-triages, builds runtime verifications, and generates scope-checked patches validated by two-sided differential tests. From 19,325 high-severity CVEs spanning 2022 to 2026, it identified 2,710 fixing commits and built 1,033 detection rules across 56 CWE families, packaged into 172 skills. Applied across 14 programs, the pipeline produced runtime evidence for 644 findings.
When LLM Decompilers Recompile More and Preserve Less
Decompilation recovers source code from binaries and underpins vulnerability detection and malware analysis, but large language model (LLM) based decompilers are now judged almost entirely by whether their output recompiles and passes shipped input/output tests, metrics that can reward code that builds cleanly yet behaves differently on other legitimate inputs or silently drops a known vulnerability. Decompile-Diverge is a behavioral comparison oracle that synthesizes a driver for each function, grows a fuzzing corpus from the reference, and reruns the decompiled code on the same inputs to detect divergence without hand-crafted tests. Across eight systems in nine configurations on established LLM decompilation corpora, candidates that pass every shipped test still diverge from the original on 4.9% of functions overall and up to 13% for one system. On 300 real GitHub library functions and 287 CVE-grounded functions, the strongest refinement LLM lifts Ghidra's build rate from 75% to 90% while its behavioral match rate falls from 74% to 62%, up to a tenth of disclosed vulnerabilities lose their crash entirely, and source-level analysis traces the divergence to invented fields, types, callees, and guards that replace the visible unknowns traditional tools leave behind.
72 more specialized papers
- Automatic Speech Recognition for Multilingual Oral History Research Sidney Wong, Chelsea Wong She, Eda Tang et al.
- EXAONE Forecast for Finance Seunghan Lee, Jaehoon Lee, Jun Seo et al.
- Low-Latency Spell Correction for Japanese Music Search Queries Anshul Garg, Pavni Tandon, Karan Bhukar et al.
- A Data Fusion Framework for Grounding Aerospace Surrogate Model via Experimental Wind-Tunnel Observations Nitin Nagesh Kulkarni, Dheeraj Vemula, Yin Yu et al.
- Corporate-Family Resolution Is Not a String-Matching Problem: A Public Benchmark Stratified by Name Visibility Harshit Gupta
- Quantum-Assisted Memory-Efficient Training for Parameter-Intensive Wi-Fi-Based Human Activity Recognition To Truong An, Jie Zhang, Guolin Yin et al.
- Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning Christos Petridis, Zoran Obradovic, Mladen Kezunovic
- BER-PEF: Unified Human Mobility Predictability Evaluation via Bayes Error Rate Estimation En Xu, Jingtao Ding, Zhiwen Yu et al.
- Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security Joshua Salako, Folajimi Osikomaiya, Olakorede Olamiju
- TNFlow: Amortized Posterior Inference for Trans-Neptunian Object Surface Composition Agastya Gaur (University of Illinois Urbana-Champaign, SETI Institute), Cristina M. Dalle Ore (Carl Sagan Center et al.
- A Constraint-Aware Generative Framework for Synthetic Origin-Destination Demand in Logistics Networks Leian Chen
- Adapting from Downturns: Prediction of Long-Term Conversational-Skill Development in Mental-Health Crisis Counselors Vivian Nguyen, Lillian Lee, Elizabeth A. Olson et al.
- Blockchain-Enabled Secure Logging for Fiscal Electronic Mechanisms: Evaluation of the Greek eSEND and myDATA Tax Systems Panagiotis Mavridis, Anargyros Baklezos, Christos Nikolopoulos
- Ultrasound-Based Prediction of Cirrhosis Decompensation Using Large-Scale Computer Vision Models Guangyi Zhang, Peiyun Ni, Eugene Cheah et al.
- Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer Fabricio C. Avini, Guilherme Trez
- Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures Amar Alem Koric, Qibang Liu, Seid Koric
- REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation Mohsen Nayebi Kerdabadi, Arya Hadizadeh Moghaddam, Dongjie Wang et al.
- Beyond a Universal Forecasting Selector: Demand-Conditioned Model Selection across Demand Patterns and Horizons Adolfo Gonz\'alez
- Recovering molecules from coarse-grained beads: free-energy-conditioned generative backmapping across chemical space Luis Itza Vazquez-Salazar, Tristan Bereau
- On-board ML for Trace Gas detection in Imaging Spectroscopy data V\'it R\r{u}\v{z}i\v{c}ka, Adam Chlus, Andrew Thorpe et al.
- A Roadmap for MEG Foundation Models Philipp Th\"olke, Hamza Abdelhedi, Yorguin Mantilla-Ramos et al.
- Uncertainty Signals for Network Intent Translation: Risk Ranking and Ambiguity Localization Ala' A. Alsamarneh, Omar Alhussein
- ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended Reality Yoga Suhas Kuruba Manjunath, Jie Gao, Lian Zhao
- BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker Seyed Mahmoud Sajjadi Mohammadabadi
- A Semantic Model of Genetic Evidence: A Step Toward Bridging the Basic-Science-Clinic Gap Michael Bouzinier, Dmitry Etin
- An Energy-Based Conservative-Dissipative Latent Neural Evolution Operator for Magnetization Dynamics Sebastian Schaffer, Lukas Exl
- Data-Driven Discovery of Composition-Dependent Constitutive Models for Hyperelasticity and Viscoelasticity of Digital Materials Josu\'e Garc\'ia-\'Avila (Department of Mechanical Engineering, Columbia University, New York City et al.
- Fast Surrogate Modeling of Excitable and Oscillatory FitzHugh-Nagumo Dynamics with Parametric Neural Operators Andrew Franck, Justin Li
- IPGeoAI: Transformer-Based Geolocation with LLM Semantic Fusion Avinash Kadimisetty, Andy Jinqing Yu, Philip Favaloro et al.
- MURAL: Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning Ahmad Mousavi (Department of Mathematics, Statistics American University), Majid Alikhani (Independent Researcher) et al.
- Too Rare to Learn: Prescribed Cyclone Tracks Degrade a Bay of Bengal Ocean Emulator Sumaiya Islam
- Continual Graph Memory for Adaptive Recommendation under Intent Drift Hao Nguyen Ngoc, Tung Nguyen, Nguyen Thi Hanh et al.
- WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding Robert Epps
- Retinal OCTA Phenotyping with LLM Reporting for Alzheimer's Disease Progga Paromita Dutta, Jeba Maliha, Md Rafiul Kabir
- Predicting Spatiotemporal Mobile Sensing-Based PM2.5 Concentrations Using Low-Rank Adapted Spatially Attentive Graph Neural Network Om Chiddarwar, Priyanka Mandal, Praveen Kumar Chandaliya et al.
- A Differentiable Neural Surrogate for Photon Propagation in Neutrino Telescopes Felix J. Yu, Berthy T. Feng, Nicholas Kamp et al.
- Wireless Foundation Models: State-of-the-Art and Open Challenges Alonso M. Pacheco Huachaca, Juan J. Rodriguez Rodriguez, Ahmed Aboulfotouh et al.
- A Fairness Audit of the Duckworth-Lewis-Stern Method: Format-Specific and Gender-Differential Bias, with an Interpretable Calibration Layer for Cricket Target Revision Soumyadeep Roy
- Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs Amrit Gopinath, Sangeetha Sivanesan
- ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing Mingrui Li, Sixian Shen, Minzhang Li et al.
- How Faithful Is Attribution for Sales Forecasting? A Counterfactual Study Glib Kechyn
- Hierarchical Possession-Aware Graph Pointer Network for Pass Receiver Selection Jingyi Wang, Da Li, Kaixin Wang et al.
- MedFlow: Class-Aware Multi-Scale Generation for Medical Time-Series Synthesis Yanhao Huang, Shibo Feng, Wanjin Feng et al.
- Long Horizon Transformer Quantile Fault Prediction for Multi Site Industrial Predictive Maintenance David J Poland, Daniele Ravi, Na Helian
- MZ-Rain: Moisture-Budget-Guided Zero-Inflated Model for Station-Level Precipitation Nowcasting Yifang Zhang, Shengwu Xiong, Henan Wang et al.
- LLM-Assisted Behavioural and Scenario Augmentation for Agent-Based Energy Adoption Models Iias Faiud, Hossein Khaleghy, Michael Schukat et al.
- Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach Iias Faiud, Jonaid Shianifar, Michael Schukat et al.
- Attention-guided super-resolution of 4D flow MRI in carotid arteries Ali Mokhtari, Dominik Obrist
- Methane Detection On Board Satellites from Unorthorectified Imagery Luca Marini, Maggie Chen, Hala Lamdouar et al.
- Physics-Aware Random Walk Fingerprints for Scalable Power Grid Graph Classification Adnan Anwar
- Qlippy: A Retrieval-Augmented GenAI Assistant for Reproducible Quantum Workflows and Experiment Tracking Mahee Gamage, Vlad Stirbu
- Beyond Co-purchase Relation: Evolution of Complementary Recommendations at Allegro Aleksandra Osowska-Kurczab, Klaudia Nazarko, Eli\v{s}ka Kosturov\'a et al.
- Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent Yunqi Zhu, Wensheng Zhang, Xuebing Yang
- NEAT-POCKET: Pocket-Conditioned Autoregressive 3D Molecular Generation with a Neighborhood-Guided Set Transformer Roxane Axel Jacob, Daniel Rose, Thierry Langer et al.
- A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment Mar\'ia Eugenia Curi, Germ\'an Capdehourat, Isabel Amigo et al.
- A Hybrid Predictive Ensemble of Machine Learning and Deep Neural Networks for Early Cardiovascular Disease Risk Assessment Balaji Venkateswaran
- SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis Yuqing Yang, Alexander Schmatz, Zhaozhao Ma et al.
- Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers Yumiao Li, Peixin Liu, Donglin Di et al.
- FedDRAW: Federated Dual Reputation Annealing Weighting for Heterogeneous Multi-Institutional Chest Radiograph Classification Maryam Moradpour, Anne-Christin Hauschild
- Hessian-based molecular conformation augmentation for a scalable and efficient strategy of machine learning interatomic potentials Bumju Kwak, Jeonghee Jo
- PRICE: A Systematic Study of LLM Adaptation Choices for Bitcoin Price Forecasting Maryam Fakhari, Mehran Safayani
- A Unified Physics-Aware Quantum Machine Learning Framework across Power GaN HEMTs and Logic Nanowire FETs: Predicting Unseen Process Splits and Held-Out Geometry Combinations with Lower Error and Tighter Split-to-Split Variability Rushat Rai, Yun-Yuan Wang, Autsada Kakaen et al.
- Self-Supervised Lexical Representation Learning for Fast, Large-Scale Phylogenetic Inference Tim Wientzek
- AI for Computational Design Science: A Responsible Human-AI Framework and Case Study on Short-Form Video Safety Surveillance Wenli Zhang, Jiaheng Xie, Zhihe Pan et al.
- Learning from VAE Errors to support ECG-based Differential Diagnosis of Myocardial Scar Shayan Sharifi, Riccardo Treu, Ilaria Gandin et al.
- LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics Gaurab Baral
- LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams Yoonju Sim, Federico Berto, Chuanbo Hua et al.
- Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education Rayed AlGhamdi
- Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation Siliang Liu, Mohammad Ghasemi, Sapan Patel et al.
- A Deep Generative Model for Synthesizing Labeled Wireless Signals Yuxiao Li, Keke Hu, Santiago Mazuelas et al.
- RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments Quoc H. Nguyen, Ali Lafzi, Abhijeet Phatak et al.
- WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data Ji Soo Lee, Xilun Chen, Pierce Chuang et al.
Large Language Models 60
Quality Recovery for Quantized KV Caches via Low-Rank Attention Adaptation
Low-bit key-value (KV) caches cut decoding memory, but the resulting quality loss varies by model and quantizer. Keeping the quantizer fixed, the method distills the floating-cache model's behavior into low-rank Q/K/V projection updates while the student runs a physically packed incremental cache. Across three seeds, 4-bit affine-cache adapters recover 54.24% of the held-out perplexity gap on TinyLlama-1.1B and 75.96% on Gemma-4-12B, and on a frozen NF4 Llama-3.1-8B base recover 60.42% under KIVI K2V2 and 37.61% under KVarN K4V2 while preserving 180-case associative retrieval; Gemma's score on a RULER subset rises from 42.80 to 48.33 against 46.15 floating. A 2-bit sweep drops TinyLlama's perplexity from 576.10 to about 11.4 versus 10.40 floating but restores only 11 to 12 of 180 retrieval cases, showing that perplexity recovery need not restore long-context retrieval.
Scalable Context Orchestration for Serving LLMs Over Voice
Voice assistants built on large language models (LLMs) typically treat conversation context as a flat, growing message list, leaving cues such as speaking rate, background noise, and packet loss implicit in the audio and leading to misaligned responses and rising costs over long sessions. llmovoice is a context-management middleware that, at each turn, assembles a bounded voice context from the current input, relevant history, and explicit paralinguistic and environmental state, then has the serving LLM reason over it to emit runtime directives that steer the response. On real voice applications and benchmarks it cuts speaking-rate alignment error by 52.4%, drops the false-interruption rate under packet loss from 46.0% to 0.9%, and lowers model usage cost by 79.2%. For long sessions it reduces per-turn cost by up to 24.9 times while keeping up to 98.7% of baseline answer quality.
Evidence Integration in Large Language Models
LLMs increasingly reason over evidence supplied by tools, retrieval, other agents, and users, yet how they fold that evidence into answers they have already started forming is poorly understood. The authors propose a distributional theory in which evidence shifts the receiver's distribution over initial answers, governed by a receiver prior weight and a candidate evidence tilt, predicting that candidates the receiver already finds probable are more persuasive, that models absorb their own characteristic errors more readily than foreign ones, and that identical evidence can help weaker models while harming stronger ones. These predictions hold across over ten million trials, twelve LLMs from four families, and eight domains including quantum mechanics, physics, genetics, and molecular biology, and models integrate candidate answers even after internally verifying them as invalid, in 93 to 100% of cases under propositional constraints. Causal interventions locate candidate integration late in the network as a structured sequence of admitting, promoting, and transporting external answers, while representations of verification are decodable but have little causal effect on the final answer.
SharedSAE: One Feature Dictionary Across Language Models
Sparse autoencoders (SAEs) are a standard tool for interpreting language model activations, but training and latent labelling are normally repeated for every model. SharedSAE pairs one shared dictionary with model-specific encoder-decoder pairs, normalizing only the selection scores so activation magnitudes are preserved, and uses model dropout so a single model can be run at inference, unlike the closest prior method, which discards magnitudes and requires all models. Trained on four 1B-scale base models from distinct families with different tokenizers, it retains 96.6% of the mean explained variance of dedicated per-model SAEs, its latents show cross-model correlations 1.8 times those of separately trained SAEs aligned post hoc, and latent descriptions transfer across models. Once the dictionary is frozen, new models can be adapted to it efficiently with near-dedicated reconstruction quality while reusing the shared descriptions.
A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models
Multilingual language models often answer semantically equivalent questions differently across languages, and methods for improving cross-lingual consistency (CLC) have been evaluated with incompatible models, tasks, and protocols. The paper runs a unified evaluation of representative inference-time and post-training CLC methods for question answering across three model families and three closed-form benchmarks, then checks on two culturally diverse QA benchmarks whether the methods suppress legitimately different answers to culture-dependent questions. Post-training methods are more reliable, with direct distribution alignment improving consistency in every model-dataset combination, while other methods are sensitive to answer format and language coverage, and cross-domain transfer is limited unless output formats match. Controlled closed-form evaluation shows no systematic loss on culture-dependent questions, but open-ended generation reveals occasional accuracy drops, especially for non-English responses.
What Attention Recalls and Recurrence Controls in Hybrid Language Models
Hybrid language models pair attention with a fixed-size recurrent state, but it has been unclear what each channel actually contributes. Two cache-level interventions probe this: split-prefill keeps only the key-value (KV) cache or only the recurrent state after prefilling a context, and state-swap pairs the KV cache from one context with the recurrent state from another in a single forward pass. On Qwen3.5 and Falcon-H1, exact retrieval survives only through attention, at 64-98% of full accuracy, and collapses to zero through recurrence, while output language and persona survive through the recurrent state and largely vanish when only the KV cache is kept. State-swap confirms the split causally, with the answer's content coming from the KV side and its language from the recurrent side, and recurrent-only generation even accepts words never seen in context that share meaning or parts with seen items.
GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion
LLMs in high-stakes settings often produce plausible but ungrounded claims, and standard retrieval-augmented generation (RAG) does little to help because it retrieves isolated passages without tracking cross-document evidence or quantifying uncertainty. GRACE breaks a model's response into atomic claims, links each to trusted knowledge priors in a weighted bipartite graph, and uses weighted centrality to classify claims as Grounded, Refuted, or Boundary, the last category capturing novel or contested claims at the edge of the model's knowledge. A Return on Attention objective sends a claim to expert review only when its priority-weighted uncertainty exceeds the cost of verification, and verified claims become new evidence anchors so the knowledge base grows over iterations. Across multiple models and both general and domain-specific datasets, the graph-grounded knowledge base outperforms RAG baselines for retrieval and the Return on Attention rule efficiently selects boundary claims worth verifying.
When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models
Expert pruning shrinks Mixture-of-Experts (MoE) models by dropping experts the router deems unimportant, but that signal collapses under over-dispersed routing, where aggressive load-balancing during training spreads tokens almost uniformly across experts. In this regime perplexity stops predicting downstream accuracy: on gpt-oss-20B the lowest-perplexity pruning configuration gives the worst mathematical reasoning while the highest-perplexity one preserves it, unlike Mixtral-8x7B-Instruct where the two degrade together. No single scoring metric wins either, since activation-aware scoring keeps math but loses 18 points on GPQA, and frequency-based scoring shows the reverse. Minimax Expert Score Allocation (MESA) iteratively boosts scores for experts serving whichever domain is currently worst hit, and at 25% expert pruning it achieves the smallest worst-case degradation across domains, beating activation-aware baselines on 7 of 11 benchmarks and generalizing to gpt-oss-120B, Gemma-4-26B-A4B, and OLMoE-1B-7B.
PerfReasoning: How Well Do LLMs Reason on Hardware Performance?
Performance modeling for hardware design and software optimization requires structured reasoning about computation, data reuse, storage, and movement. PerfReasoning evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code: given workload, architecture, and mapping specifications, models must compare mappings and predict off-chip traffic and buffer requirements. The strongest closed models exceed 90% on reasoning question answering and the best open-weight model reaches 82.4%, but model construction is much harder, with only GPT-5.6 Sol exceeding an 80% pass rate while all other configurations average below 15% and vary widely across runs. Task-specific reinforcement learning lifts a 4B model's mapping-reasoning accuracy by 15.7 points, whereas feedback-free multi-round self-revision prompting is not reliably effective.
Scale-QLoRA: Code-Invariant Adapter Merging for Native 4-bit Microscaling LLMs
Merging a LoRA adapter into its base model is standard deployment practice, but on native 4-bit microscaling checkpoints such as NVFP4 and MXFP4 the merged weights must pass back through a quantizer that re-derives the discrete E2M1 code plane, roughly 90% of the artifact's bytes, coupling the result to one quantization convention and, done naively, deleting the adaptation by up to 39 percentage points because the reconstruction optimum against an already-on-grid base is the base itself. Scale-QLoRA instead trains only the native per-block scale field on the deployment grid and freezes every E2M1 code, so within a fixed format, scale grid, and block layout the merge becomes a bit-exact identity and the artifact is code-invariant. Across four models and four tasks it is accuracy-lossless, matching merge-aware QAT-LoRA, while sidestepping quantizer sensitivity where rounding-rule mismatches can drive weight-space artifacts toward 0%. Freezing codes also removes the weight-space straight-through estimator from training, giving 3.9x faster steps on a dense 8B model, and enables exact rollback, code-plane deduplication, and roughly 125x faster scale-only task swaps.
Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One
Language generation is nearly always sequential, either token by token in autoregressive models or through long iterative refinement trajectories in diffusion language models. PlaidQ is a 0.7B continuous diffusion language model for code that repurposes a pretrained autoregressive model as a bidirectional denoiser over continuous token embeddings, then distills its trajectory with distribution matching for few-step generation and paired-trajectory supervision for one-step generation. At matched scale it is competitive with discrete diffusion models on code, and a 16-step distilled student reaches 31.78 and 40.49 pass@10 on HumanEval and MBPP+, beating its own teacher sampled for 512 steps. A single denoising step still yields functionally correct programs at 7.07 pass@1 on HumanEval, and code and checkpoints are released.
From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs
Reliable LLM deployment requires distinguishing uncertainty that stems from irreducible task ambiguity from gaps in the model's knowledge. Existing decomposition methods generate multiple clarifications of an ambiguous input, answer under each, and compare the answers, but the authors argue theoretically that the answers are redundant, costly, and prone to epistemic leakage, and instead estimate ambiguity-induced aleatoric uncertainty directly from the space of plausible interpretations. On ambiguity detection across three benchmarks, the clarification-only approach raises AUROC to 63.34 from 60.85 while cutting output tokens 4-26x and API calls 2.2-3.5x, and its estimates correlate substantially less with epistemic uncertainty.
Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models
Fine-grained Mixture-of-Experts (MoE) models route each token to the top-k experts and renormalize router probabilities, and the authors show this renormalization implicitly calibrates expert output gain to the training-time k, so reducing k at inference changes not just which experts fire but the strength of the expert branch. Their fix activates the top k1 experts while normalizing by the probability mass of the top k2 experts, adding a single integer with no parameters, training, or measurable compute overhead. On Qwen3.6-35B-A3B, going from 8 to 4 experts costs 4.65 MMLU points under standard renormalization but only 0.35 points with k2=16 while halving routed-expert compute, and the result replicates on the 11x larger Qwen3.5-397B-A17B, where 10 to 5 experts loses only 0.55 points. Removing renormalization entirely is catastrophic, expert identity matters far more than weighting, and perplexity and downstream accuracy favor different k2, so compression settings should not be chosen from unlabeled text alone.
Optimizer Memory Schedules for Outscaling the Overtraining Axis
Optimizers are usually compared at a single training horizon, yet relative performance and optimal hyperparameters can shift substantially with the amount of overtraining. The authors compare the matrix-preconditioned Muon and SOAP, the momentum-scheduled ADANA, and AdamW on models from 51M to 253M parameters across overtraining factors from 1x to 256x, sweeping base learning rate at every setting. They find the preferred learning rate schedule can reverse across the overtraining axis, the best weight decay scales roughly as the square root of the overtraining factor, and longer horizons favor longer fixed memory. With log-time weight decay and momentum cooldown, ADANA outscales AdamW with an exponent advantage close to that predicted by DANA theory, and while Muon and SOAP hold roughly constant token-efficiency advantages over AdamW, ADANA overtakes Muon and becomes competitive with SOAP at the highest overtraining factors.
When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models
A 0.6B language model asked to verify 1,200 logical conclusions, half corrupted by a single semantic edit, answers YES every time, yet linear probes on its hidden states read the correct verdict at 0.96 AUC and transfer to unseen logical structures. The authors trace where the verdict is lost and find it survives to the output logits, where the margin still carries 0.89 AUC along a well-aligned readout direction, so a single saturated decision threshold offset by +4.6 sigma is what erases the behavior. Across 90 semantic-label configurations spanning five models and three families, behavioral accuracy collapses onto a function of threshold offset (Spearman -0.93), a one-parameter correction never fit on evaluated structures repairs accuracy from 50% to 81% at 0.6B, calibrated margin decoding recovers 94% at 8B, and few-shot prompting is shown to work by the same recentering mechanism. Comparing probe to margin separates concealed, miscalibrated, and undetected regimes, and the authors caution that in standard generation settings, answer-surface features and heuristic labels can reproduce published probing results without any internal access.
Choosing the Right Language Mode at Inference Time for Multilingual Reliability
Multilingual large language models reason poorly in low- and mid-resource languages, and while translating into English can help by tapping stronger English-centric representations, it is unclear how much translation helps before it starts causing interference and overconfidence. Experiments with LLaMA and Qwen models vary text scope and language mode (target-only, English-only, bilingual) and find a trade-off: English context improves understanding and recovers errors caused by non-English comprehension, but redundant bilingual context intensifies interference. Reliability-Aware Adaptive Inference (RAAI) is a training-free test-time method that routes and fuses prompts based on Expected Calibration Error (ECE) and uses a mid-layer Risk Index (RI) to gate sequential reasoning, spending extra compute only where it is likely to help. Across the two model families, RAAI raises accuracy by 25 to 37.7% on low-resource languages and lowers calibration error by roughly 3 to 6%, with the largest gains in the lowest-resource language tiers.
Model Retirement Creates Reproducibility Risk in Biomedical AI Publications
Commercial large language models (LLMs) are retired on deprecation schedules, which threatens the computational reproducibility of research built on them. The authors searched PubMed for original research articles from 2022 through March 2026 that applied a specific LLM to a biomedical task, used an extraction agent to pull model names from 61,077 abstracts with human validation on a subset, and compiled release and retirement data for the 50 most frequently used models. Across 8,931 paper-model mentions in 5,242 publications, 77.7% cited a closed-weight commercial model and 42% involved a model already retired at publication or scheduled to retire within two years, with a median of 538 days from publication to retirement.
Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving
Prefix caching, which reuses the key and value tensors of a shared prompt prefix across requests, is on by default in major open-source serving stacks and assumed to be a transparent optimization. Holding model, decoding parameters, seed, and request order fixed and issuing every request serially at batch size one, the authors ran an eighty-episode multi-turn agentic tool-use workload with caching enabled and disabled across two engines and four weight formats. Enabling the cache changed the agent's trajectory on 36.2% of episodes at 16-bit precision and on 75.0% at four-bit, while cache-disabled repeated runs were bit-identical in all 800 episodes; follow-up experiments trace run-to-run divergence to a single server-level prompt-cache setting and show that cached serving is deterministic given cache state but irreproducible in practice because that state is absent from the request and never reset by default.
Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges
Diffusion language models (DLMs) generate text by iteratively denoising tokens rather than decoding left to right, which lets them refine many uncertain tokens in parallel, use bidirectional context, and trade quality for latency more flexibly than autoregressive large language models (LLMs). This survey argues these properties suit agents running on mobile edge devices, where partial refinement, early exit, and constraint-guided correction can cut response delay and communication cost under noisy or incomplete context. It reviews DLM foundations and maps them to edge constraints on latency, memory, energy, bandwidth, privacy, and reliability, covering efficient architectures, training and inference acceleration, compression, edge/cloud deployment, Internet of Things (IoT) and wireless applications, and agent evaluation, before listing open problems in long-context state management, split inference, and reproducible benchmarking.
Can Activation Steering Capture Multidimensional Authorship Style?
Activation steering works for well-defined attributes, but authorship style is multidimensional and hard to specify in words. The authors construct per-aspect steering directions from structured contrastive prompts along rhetorically motivated dimensions and find that the directions share a common authorship backbone while conflicting on aspect-specific residuals, which explains why naive averaging of directions fails. Aspect-Aware Activation Steering (A3S) merges the per-aspect directions with interference-aware aggregation and tunes steering strength per instance, improving style transfer where styles are genuinely multi-aspect, beating a trained baseline in preference evaluations on out-of-domain benchmarks, and keeping overlap with target exemplars low.
When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models
Financial large language models used to summarize reports often fabricate numbers, and prior work has blamed weak numerical reasoning without testing that assumption under controlled fine-tuning. The authors compare a base instruction-tuned model against a domain language-adapted variant (FT-A) and a numeracy-enhanced domain variant (FT-A+B+C), scoring outputs with a three-level taxonomy that separates overt currency fabrication, covert-explicit professional-convention numbers, and covert-implicit ungrounded quantitative claims. Overt hallucination rises from 5.4% for the base model to 82.5% after domain adaptation and 98% after adding numeracy supervision, with the fine-tuned models frequently injecting memorized template values regardless of the input. The authors conclude that domain adaptation erodes numerical restraint rather than reasoning ability and recommend evaluating all detectability levels plus grounding-aware generation or abstention at deployment.
A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures
Multilingual language models develop shared cross-lingual representations, and several interpretability metrics claim to quantify that sharing, but they were developed in isolation and their disagreements are hard to interpret. The authors compare four metrics, CKA, ANC, Gaussian mixture model (GMM) dominance per token, and ILO, across 21 base models from five families spanning 125M to 14B parameters, correlating each with cross-lingual transfer on five downstream tasks. The metrics disagree substantially, and the authors trace the disagreement to anisotropy, the tendency of representations to cluster in a narrow cone of embedding space. Only ILO retains a strong correlation with cross-lingual transfer (Spearman's rho of 0.90) after controlling for model size, family, and task, so they recommend it as the primary sharing metric, reported alongside anisotropy diagnostics.
On Epistemic Diversity in Large Language Models
Large language models increasingly answer questions, explain, and teach, where a correct answer can still narrow a user's access to alternative valid answers and reasoning routes. Drawing on philosophy and social epistemology, the authors formalize epistemic diversity for LLMs as the range of valid answers, explanations, and reasoning paths a model exposes, argue it is a needed evaluation dimension for knowledge-intensive use, and propose a preliminary framework for measuring it, operationalized in two domains. Frontier LLMs frequently exhibit epistemic narrowness, collapsing large spaces of valid answers onto small canonical subsets. The authors conclude that evaluation should move beyond accuracy-oriented paradigms and treat epistemic diversity as a model capability.
Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference
Mixture-of-Experts (MoE) models activate only a few experts per token, but the full expert set often exceeds GPU memory, so decoding repeatedly transfers weights. This work treats expert-cache management as a model-side algorithmic problem, jointly post-training the backbone with lightweight auxiliary cache routers while preserving the native top-K selection rule at inference: a Temporal Router predicts same-layer reuse and retains experts without proactive loading, and the full Spatio-Temporal Router additionally uses the causal predecessor's hidden state to refine the cache before the target layer is accessed. Evaluated on Qwen3 and GPT-OSS across GSM8K, MATH, and CommonsenseQA, the full router on Qwen3 improves load-adjusted hit rate by 1.15 to 18.03 points and reduces expert-weight traffic by 4.6 to 53.3% versus the strongest prefetching baseline, with competitive but task-dependent results on GPT-OSS. An auxiliary-only ablation keeps baseline accuracy but yields far smaller cache gains, showing the joint post-training does the work.
Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair
Evaluations of large language models (LLMs) for automated program repair (APR) mostly score final patches and say little about where the model's reasoning departs from the actual repair evidence. The authors define hallucination as producing patches or intermediate artifacts not grounded in that evidence, and measure it both in final patches and in three understanding tasks: identifying the triggering test case, predicting line coverage, and generating additional test cases. Running three LLMs on 832 Defects4J bugs, they find only 21.0% to 55.9% of patches pass the developer test suite, and manual analysis of 812 sampled repairs flags repair hallucinations in 72.7% of cases, including patches that pass every available test, with wrong causal localization (45.9%) and wrong repair strategy (18.5%) as the leading causes. Better intermediate artifacts generally accompany successful repairs, but the correlation is not reliable, and models often misidentify triggering tests, mispredict coverage around branches, and write tests that miss the bug-triggering condition.
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
Large Reasoning Models (LRMs) generate long chains of thought whose key-value (KV) cache grows linearly and often exceeds GPU memory, while existing compression methods pick tokens to keep based on recent queries, assuming those predict future attention. The authors show this assumption breaks in long reasoning: certain Thought Revisiting Tokens re-attend to distant context such as plans formed early in the trace, and the queries behind them fall into a small number of clusters in embedding space. BeaconKV is a training-free method that keeps compact beacon queries representing each cluster to anticipate which KV pairs will be revisited without storing the full query history. Across four open-source LRMs and diverse reasoning benchmarks it generally outperforms existing compression methods, achieving up to 5.8x memory reduction while nearly matching full-cache accuracy and improving throughput by over 4.3x.
A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering
Structured retrieval-augmented generation (RAG) methods that reason over trees or graphs help with multi-hop questions but struggle with evidence-intensive QA, where answers must synthesize information spread across dozens or hundreds of documents, because their structures are rigid and evidence is gathered without regard to the reasoning topology. APT-RAG expands its reasoning structure adaptively based on question dependencies and evidence needs, and gathers evidence in a topology-aware way through sibling evidence reuse, direct retrieval, and aggregation from child nodes, with evidence-guided batched answer generation to cut generation overhead. On evidence-intensive QA benchmarks it outperforms existing structured RAG methods, and code is available.
Amortizing Scaling Law Construction Costs
Fitting a scaling law needs only the best-loss frontier across compute scales, yet standard practice trains an exhaustive grid over hyperparameters, token budgets, and parameter counts and then discards most of it. The proposed framework casts data collection for scaling laws as a Bayesian optimization problem and introduces metrics for comparing fitting methods under constrained compute budgets. Progressively expanding the compute budget during acquisition, mirroring the compute-ordered evaluation of configurations in practice, substantially improves recovery efficiency, and augmenting the observed runs with surrogate-fantasized evaluations reconstructs the broader experimental grid. Together these closely match dense-grid scaling law fits at 10 to 100 times lower compute cost.
Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection
Detecting hallucinations, outputs that are factually incorrect or unsupported by the source, is hard because large language models give little insight into why a response may be inaccurate. The approach tests whether a low-level symbolic competence such as SQL can serve as an unsupervised grounding mechanism for a high-level task: an LLM first builds an SQL database from the reference documents, and the detection pipeline then reasons over both the reference and the sampled response through that database, providing a neurosymbolic check. On the RAGTruth and DiaHalu hallucination detection datasets, the method improves on direct prediction and competes with state-of-the-art detectors without any domain-specific fine-tuning, relying only on a general competence already present in LLMs.
EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages
Machine translation (MT) is a scalable way to extend English instruction-tuning data to other languages, but it can distort task-critical constraints and required outputs, producing corrupted training examples that degrade the models trained on them. EuroAlpaca is a task-preserving localisation pipeline and near-parallel resource covering 50 European languages and regional varieties that either applies field-wise MT while preserving task-critical content or reconstructs a task-equivalent target-language instance, followed by validation of cross-field coherence and target-language consistency, and it is paired with European-IFEval, a multilingual benchmark for verifiable instruction following. In LoRA experiments with four LLMs, directly translated data improves ROUGE-L and F-BERT on the Aya Evaluation Suite but cuts European-IFEval accuracy by 29.8% relative to the unadapted baseline, whereas adaptation with EuroAlpaca raises accuracy by 12.9% over the same baseline while also achieving the highest ROUGE-L and F-BERT scores on Aya.
Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time
Attention-head attribution in Transformer classifiers usually forces a choice between fine-grained circuit tracing and coarse output-level probing. The authors define an influence score that combines each head's directional effect on the logits with its structural contribution to the residual stream, so it can be aggregated at the head, layer, and whole-network level. Applied to a DeBERTa model fine-tuned for prompt injection detection, the score exposes distinct head-level decision patterns for correct versus erroneous predictions, offering a middle ground between circuit analysis and global output-based interpretability methods.
Single-Query Black-Box Calibration Auditing via Logit Bias
Standard calibration metrics need the continuous output probabilities that commercial LLM APIs increasingly withhold. The authors show that any API exposing a logit_bias parameter can be manipulated to test exact probability thresholds with strictly one query per sample, and build on this a provably consistent estimator of the True Calibration Error for binary classification tasks. The result is an efficient method for auditing the calibration of black-box foundation models used as zero-shot classifiers.
Large Language Models with At Most One Spike per Neuron
Spiking neural networks (SNNs) promise energy-efficient large language models (LLMs) through sparse, event-driven computation, and time-to-first-spike (TTFS) coding pushes firing rates to at most one spike per neuron per time window. Conventional TTFS networks cannot express operations such as layer normalization and matrix multiplication, so the authors introduce a reference-based encoding strategy for embedding layers, layer normalization, attention-related operations, and dropout, then build and train a fully TTFS-based architecture end to end. On BERT and GPT-2, the spiking models match their artificial neural network counterparts on natural language understanding and common-sense reasoning, while a clear gap remains on language-modeling perplexity; the authors describe this as the first TTFS spiking LLM scaled to 1.5 billion parameters. Reported energy figures are a spike-count proxy under an established cost model rather than measurements on neuromorphic hardware.
Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG
Retrieval-Augmented Generation (RAG) improves answers with retrieved documents, but long contexts slow inference, and soft compression methods that shrink each document into a short embedding sequence are usually trained by distilling from the uncompressed system, capping them at its performance. DEX-Comp uses a two-stage recipe: Pure Distillation warm-starts the compressor only on responses the uncompressed RAG got right, then Hard Exploration runs reinforcement learning solely on queries the uncompressed system fails, pushing the model toward computation patterns better suited to compressed inputs. Across five open-domain QA benchmarks at retrieval depths from top-5 to top-30, DEX-Comp compresses contexts 16x and speeds inference 4x to 24x while matching or exceeding the uncompressed baseline. Ablations across datasets and backbones attribute distinct contributions to each stage.
What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection
On-Policy Distillation (OPD) is a common post-training method for reasoning models, but which training data actually drives its gains has been little studied. The authors first try 1-shot OPD, training on a single example, and find it consistently effective across sampled examples, with harder problems giving larger gains; analysis shows the improvement comes not from high token entropy but from the longer chain-of-thought (CoT) paths that hard problems produce, which keep the student aligned with the teacher over long horizons and teach patterns such as reflection that short CoTs lack. They propose selecting only hard examples, including ones that completely exceed the teacher's own ability. Across four models from 1.5B to 7B parameters, training on just 8 selected hard examples matches a 17K-example baseline.
ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs
Mixture-of-Experts (MoE) language models route every token to a fixed top-k set of experts, spending compute on experts that contribute little, and existing skipping methods depend on router confidence, calibration data, or extra training. ACE is a training-free, calibration-free scheme that scores each routed expert with two offline-computed views, a Global Spectral Proxy (GSP) derived from the gate, up, and down projections with RMSNorm scaling, and a Router-Conditioned Refinement (RCR) that measures expert responses along routing-preferred directions, and skips a slot only when both views and the runtime gate agree it is low-contribution, always keeping the top-1 expert. Across three MoE models and eight benchmarks it beats static and dynamic baselines with the gap widening under aggressive skipping; at 50% expert skipping on Qwen3.6-35B-A3B it lowers WikiText-2 perplexity by 7.96% and raises average downstream accuracy by 4.15 points over the strongest competitor.
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Layer dropout, also called stochastic depth, speeds training and enables zero-shot layer pruning in vision and language transformers, yet it has vanished from LLM pretraining recipes amid reports that it hurts accuracy, without any systematic study of the effect. Across more than 2,400 pretraining runs on Cerebras CS-3 systems spanning 271M to 8.2B parameters and up to 160B tokens, the authors establish best practices for the layer distribution, time schedule, and optimizer hyperparameters and show that at equal training FLOPs layer dropout yields lower loss. Models reach lower or similar validation loss while saving up to 25% of training FLOPs, and the trained models support post-training optimizations such as early exit, intermediate-layer skipping, and self-speculative decoding for up to 1.5x inference speedup with negligible accuracy loss.
Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods
Meta-evaluation of reference-based automatic metrics for natural language generation usually measures agreement with human judgments, which says little about how an evaluator behaves under controlled conditions. The authors propose behavioral correctness assumptions: a taxonomy of correctness-preserving and correctness-altering response transformations, each paired with the scoring behavior a sound evaluator should exhibit. Applying this to lexical, character-level, semantic, LLM-based, and hybrid evaluators, they analyze assumption-level behavior, stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility. No evaluator satisfies all the proposed assumptions, and evaluators with similar aggregate scores can have substantially different behavioral profiles.
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
On-policy distillation (OPD) gives dense per-token supervision for post-training language models, but external teachers suffer from distribution mismatch and self-distillation with privileged context is limited by in-context learning capacity. RISE (Recursive Improvement via Self-Extrapolating Policy Distillation) instead builds a synthetic teacher from the model's own reinforcement learning with verifiable rewards (RLVR) trajectory, extrapolating the displacement between the current checkpoint and a trailing anchor in parameter or logit space to turn a sparse outcome-driven update into a dense token-level target. Because the teacher is refreshed every iteration as the student improves, distillation becomes a recursive loop in which outcome rewards ground the extrapolation and the extrapolated teacher refines token-level decisions. Across mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks, RISE outperforms both RLVR-only training and on-policy self-distillation.
How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing
Hyper-Connections and their manifold-constrained variant mHC widen the residual pathway from one stream to several, but it has been unclear how trained models actually use this extra capacity. The authors analyze the four-stream residual pathway of DeepSeek-V4-Flash using effective stream counts, cross-stream residual weights, and inter-stream cosine similarity. A typical attention or feed-forward site effectively reads and writes about two streams, the dominant stream changes across depth, and residual mixing is concentrated in early layers, with layers 22 through 42 mostly carrying each stream forward separately. Replacing the late mixers with identity raises C4 perplexity by only 1.9% and preserves the six-task average, while replacing the early mixers raises perplexity by 41%, indicating the model realizes only part of the flexibility mHC affords.
Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
Accuracy on molecular property benchmarks cannot distinguish a large language model (LLM) that predicts a property from one that recalls a published number. The authors audit 22 frontier models on 12 regression benchmarks for digit-level verbatim retrieval and find it widespread but benchmark-specific: on five datasets more than 50% of the models show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. Running the same prompts on the same molecules at a higher reasoning level flags retrieval 89% more often than at the lowest level. An attempt to interrupt retrieval in the most contaminated cases shows the strongest models sometimes still recognise transformed SMILES strings paired with original labels, and suppressing retrieval pulls the models' relative prediction errors closer together, indicating that general predictive ability is not determined solely by how many values a model has memorised.
19 more specialized papers
- You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments Shiwei Hong, Junjie Ma, Emma Jiren Wang et al.
- Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation Giulia Pucci, Ruizhe Li, Arabella Sinclair
- LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar et al.
- A Calibrated Reflection Approach for Enhancing Confidence Estimation in LLMs Umesh Bodhwani, Yuan Ling, Shujing Dong et al.
- Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines Siddharth Vohra, Runmin Jiang, Xiaomo Li et al.
- PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning Taegyun Kim, Youngwook Ham, Jungwook Rhim et al.
- CAGE: Coherence-Aware Graph Encoding for Retrieval-Augmented Generation Tong Qi, Jingyu Wu, Youbing Yin et al.
- How Do Language Models Represent and Use Phonological Information for Allomorph Selection? Sangwoo Kim, Sangah Lee
- PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces Xinyu Li, Hao Zhou, Jianfeng Zhu et al.
- Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM Xinyu Li, Ruoming Jin, Jianfeng Zhu et al.
- Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3 Giang Son Nguyen, Nhi Ngoc-Yen Nguyen, Wray Buntine et al.
- Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities Arnau Ayguad\'e Domingo, Stefan Bott, Horacio Saggion
- Generating Constructive Feedback on Stories via Reinforcement Learning Maja Stahl, Timon Ziegenbein, Henning Wachsmuth
- CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation Suhyun Lee, Wenxuan Zhang, W. Quin Yow et al.
- Discourse Dependency: A Continuous Criterion for Translation Difficulty Ahrii Kim, Chanjun Park, Seong-heum Kim
- MoirfEolas and Cr\'iochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology Jane Adkins, Abigail Walsh, Brian Davis et al.
- Repeated Queries Exhaust an LLM's Brand Recommendations but Not Its Sources Dmitrij \.Zatuchin
- NS-ST-GraphRAG: Neuro-Symbolic Spatio-Temporal GraphRAG for Literary Knowledge Processing Zheng Kui Lin
- Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models Jos\'e Luciano Ver\c{c}osa Marques, Frederico Jorge Heitmann, Daniel Omar Perez et al.
Other 51
Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters
Distributed training repeatedly exchanges data among GPU nodes, so congestion on a single flow can stall an entire communication round, and existing remedies assume global control over all jobs or switch-level support that a single user in a shared cloud cannot rely on. REACT operates at the communication-library layer as a shim over NCCL, detecting congestion at runtime from readily available flow statistics and retuning the collective pattern, for example by changing which node aggregates in an AllReduce tree, while preserving the semantics of the exchange. On a shared academic GPU cluster it improves algorithm bandwidth by 13% to 38% under network congestion, and simulations across congestion scenarios show gains of up to 75%, with no support required from the underlying network infrastructure.
When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
In quantized recurrent networks, the stored low-precision state is fed back at the next time step, so the rule used to write that state can change all later computation. The authors name this rule recurrent-state write-back and isolate its effect in a compact GRU encoder-decoder for fluorescence lifetime imaging, which must estimate two lifetime parameters from very noisy time-resolved signals. With the trained model held fixed, switching to deterministic 4-bit state storage raises estimation errors by roughly 70x and 300x for the two parameters, because repeated small updates fall below the write threshold and the stored state freezes while the network keeps proposing change. Error feedback, residual memory, and direction memory carry the suppressed updates across time and restore accuracy without retraining, higher state precision can worsen a fixed solution, matched training can learn compatibility with the state interface, and an independently trained LSTM reproduces the failure with the cell state more sensitive than the hidden state.
Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters
Compound AI workflows chain multiple models and software stages, and deploying one on a heterogeneous cluster means choosing a model variant per stage and a placement that satisfies service-level objectives (SLOs). System metrics can be profiled per stage and composed, but accuracy cannot, since upstream errors propagate downstream; end-to-end profiling scales poorly and product-of-stages surrogates misrank candidate plans. Atlas introduces MAP, a Markovian Accuracy Predictor that buckets intermediate outputs and composes local conditional accuracy transitions between adjacent stages along the workflow topology, then solves plan selection as a mixed-integer linear program maximizing predicted accuracy under SLOs. Across four workflows, MAP reaches Spearman correlation up to 0.947 with 2.6x less profiling, and the optimizer picks plans within 0.03 of oracle accuracy while cutting deployment cost by up to 42% through heterogeneous placement.
Mitra-v2 Technical Report
Mitra-v2 is a tabular foundation model for real-world classification and regression, trained entirely on synthetic data from a pretraining distribution much larger and more diverse than Mitra-v1's, with a small 2D Transformer backbone that supports longer contexts and larger feature spaces. On the TabArena and TALENT benchmarks spanning more than 300 real datasets, it performs at the level of the industry-scale TabFM and EXAONE Tabular models and clearly outperforms TabPFN-3 and TabICLv2. It matches the 1.6B-parameter TabFM with only 77M parameters, about 5% of the size, and ranks first on classification tasks with more than ten classes despite pretraining only on tasks with at most ten. Weights, inference and fine-tuning code, and evaluation results are released under Apache-2.0.
SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds
Generative audio models for everyday sounds have grown so large that synthesizing them demands industrial-scale compute and massive datasets. SCAPES is a lightweight model that synthesizes environmental sound textures under high-level semantic control by operating on the continuous latent space of a neural audio codec rather than discrete tokens, splitting audio into overlapping segments and modeling the evolution of latent trajectories with a Continuous Normalizing Flow (CNF) trained via Flow Matching. A 36-million-parameter instance trains on small, uncurated datasets using a single consumer-grade GPU, converging after a training time of roughly twice the duration of the source audio, while producing outputs with long-term stability, semantic consistency, and smooth interpolation between semantic conditions. Code, pretrained weights, audio examples, and an interactive demo are public.
Dynamic Heterogeneous Graph Representation Learning: A Survey
Real-world networks combine multiple node and edge types with structure that changes over time, and static or homogeneous graph representation learning (GRL) handles neither well. This survey gives the first systematic review of representation learning for Dynamic Heterogeneous Graphs (DHGs), starting from a unified formal definition that covers both discrete-time and continuous-time graphs. It organizes the literature into an algorithm-centric taxonomy spanning early embedding methods, graph neural network (GNN) models, and recent Transformer-based approaches, highlights each family's modeling bias with respect to temporal granularity, and summarizes applications, datasets, benchmarks, and open directions.
From Deep to Shallow: Unconstrained and Efficient Layer Merging Strategy
Depth compression shrinks networks by finding redundant activation functions, linearizing them, and folding the surrounding layers together, but existing methods cannot handle convolutions with padding because no analytical merge solution exists, and they enlarge the kernels of merged layers, which eats the speed-up. The strategy proposed here merges layers that lack an analytical solution and does so without any kernel-size growth. Validation covers multiple architectures and datasets, with inference speed-ups measured on real embedded platforms and code released publicly.
Fast Gauss Sums via Flash Attention
Weighted sums of Gaussian kernels underpin maximum mean discrepancies (MMDs), kernel gradient flows, Stein variational gradient descent (SVGD), and many other kernel methods, yet they lack the heavy hardware-level optimization that softmax attention has received. The authors show that two small augmentations of the inputs turn the normalized softmax reduction computed by flash attention into an unnormalized Gauss sum with arbitrary signed weights, so existing attention kernels evaluate it without any custom GPU code. For feature dimension above 8 in fp16, the approach beats both compiled PyTorch and PyKeOps kernels in speed, memory overhead, and accuracy, and memory usage stays linear in the number of points.
43 more specialized papers
- How Much Does Corpus Choice Change Dependency-Distance Estimates? Sirui Chen
- Self-Supervised Pretraining of Molecular Graph Encoders with LeJEPA Micha{\l} Kulczykowski, Rafa{\l} {\L}ab\k{e}dzki
- ProToMEx: Rapid, Interpretable Explanations via Structured Representations Athina Georgara, Adarsh Valoor, Sarvapali D. Ramchurn
- Compute-in-Memory Attention: A Time-Domain Analog Softmax Circuit with RC-Tunable Temperature Ankur Singh, Ashish Gautam, Shruti R. Kulkarni et al.
- Memory as transformation: LETHE, a self-referential gan-inspired architecture Francesco Vitucci, Anthony Di Furia, Francesco Scagliola
- A Quantum Variational Approach to Prototypical Recurrent Unit Mahyar Sadeghi Garjan, Tommaso Cesari, Michel Barbeau
- Evaluation of Phonetic Encoding Algorithms on Transcription Datasets Can \"Ozbey, Emre Kaplan, Berkin Deniz Kahya
- The Anatomy of an ASR Hallucination Hamees Sayed, Apoorv Singh, Kumar Aman et al.
- Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes Yushi Ye, Wilson Zheng, Yongyi Zang
- A Sim-to-Real Study of Surface-Code Decoder Benchmarking Shay J. Manor, Leila S. Erhili, Yassine Jebbouri
- JLIR: A Julia-Native MLIR-Inspired Intermediate Representation with Automatic JACC Kernel Extraction Narasinga Rao Miniskar, Seyong Lee, Keita Teranishi et al.
- Hidden In Plain Gaze: Gaze Representations as Privacy Controls for Utility and Re-identification Risk in XR Cory Ilo, Brendan-David John, Doug A. Bowman
- GNN-Guided Graph Coarsening and Adaptive QUBO Penalties for the Capacitated Vehicle Routing Problem with Time Windows on a Quantum Annealer Youssef Kamel Rezk, Pawe{\l} Gora
- SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery Mansooreh Montazerin, Antonio Ortega, Ajitesh Srivastava
- Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning Ming Xiang, Stratis Ioannidis, Edmund Yeh et al.
- When Does an Interpretation Count as Established? The Formation, Evaluation, and Responsibility of Interpretation in Generative AI Deyu Jing
- A Robust Watermark-based Fingerprint Framework for GNNs Ownership Verification Han Zhang, Yan Wang, Guanfeng Liu et al.
- Communication-Efficient Personalized Federated Learning via Layer-Wise Multi-Threshold Random Sketching Xu Zhang, Xingyu Hou, Jiacheng Cheng et al.
- PACE: Propagation-Aware Collaborative Correction for One-Shot Personalized Federated Graph Learning Ruizhe Huang, Chengran Li, Xiaochuan Shi
- MARLA: A Conceptual Scaffold for Regulatory Learning under the EU AI Act Alessio Buscemi, Tom Deckenbrunnen, Imane Hmiddou et al.
- TreeFI: Value-Aware Statistical Fault Injection for Deep Neural Networks Noam Bires, Marcello Traiola, Angeliki Kritikakou et al.
- Why We Care About Understanding: Competence through Predictive Compression Matthieu Queloz, Pierre Beckmann
- Global to Local: Topology-Preserving Adaptive Graph Pooling via Granular-Ball Sen Zhao, Gaojie Xu, Shuyin Xia et al.
- Beyond Homoscedasticity: Decoupled Uncertainty Optimization for Deep Imbalanced Regression Juncheng Zhou, Jiaxi Lu, Weijing Zeng et al.
- Solution-space heterogeneity shapes federated learning dynamics across partial differential equations Ping Luo, Jiahuan Wang, Ziqing Wen et al.
- Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications Yanchen Li, Xiaoming Xue, Kay Chen Tan
- Impact of Data Loss in Postprocessing on Training and Inference of Quantum Neural Networks Soraya V. Panambalom, Edoardo Altamura, Nick Chancellor et al.
- Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless Networks Abdessamed Qchohi, Jessica Moysen Cortes, Matteo Zecchin
- MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning Guanglong Sun, Kanglei Zhou, Liyuan Wang et al.
- ProCA: Progressive Contrastive Alignment for Robust EEG Visual Decoding Kanglei Zhou, Chunyan Lan, Dongyang Li et al.
- Improving Language Identification for Code-Switched Utterances with Integer Linear Programming Joanna Rado{\l}a, Josep Maria Crego, Fran\c{c}ois Yvon
- A Comparative Study of Counterfactual Explainers for Graph Neural Networks Enabling Multiple Types of Graph Edit Maria Myrto Villia, Filippos Gouidis, Theodore Patkos et al.
- MomentQuant: an even more minimalist interval method with linear time complexity for time series classification Johann Faouzi
- From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned Baseline Andrew James Amos
- Beyond Stationarity in Time Series: Discovering Causal Structures and Latent Regimes via Markov Blankets Lei Zan, Charles K. Assaad, Emilie Devijver et al.
- Measuring the Novelty of Biomedical Papers Using the Latent Distances between Knowledge Units Yi Zhao, Heng Zhang, Yuzhuo Wang et al.
- Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets Arunan J
- FluxDisco: Symbolic Regression for Stoichiometric Dynamical Systems via Monte Carlo Graph Search Cassandra Durr (Lancaster University), Alvaro K\"ohn-Luque (University of Oslo), Chris Jewell (Lancaster University) et al.
- Proton Irradiation Characterization of an Open-Source ML Accelerator on a Zynq UltraScale+ MPSoC Saad Memon, Rafal Graczyk, Jan Swako\'n et al.
- GLASS: Graph-Language Alignment with Spherical Scoring for Transferable Graph-Level Anomaly Detection Xudong Wang, Chris Ding, Tongxin Li et al.
- Embedded Graph Flows for Categorical Graph Generation Ethan Ma, Zihan Wang, Chris Siu Yeung Chow et al.
- Variational Continuation for Double Pendulum Periodic Orbits Leo Yao, Ziming Liu, Max Tegmark
- Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction Sihwa Park
Agents 50
Interface-Induced Trajectory Censoring
Agent benchmarks read tool-call rates off the serving stack, but that number can be zero while the model is emitting well-formed calls, because the interface drops them before the executor or scorer sees them. On BFCL v4, holding weights, cases, decoding, and seeds fixed and changing only the serving adapter, the same model scores 0.00 or 0.96, and a 2x2 over chat template and parser puts the entire effect in their interaction, so fixing either component alone buys nothing. The same swap on tau-bench moves server-parsed calls from 0 to 636, and inside verl's AgentLoop at 7B, 45 of 115 generations carry a complete call yet none are accepted, executed, or returned as an observation; repairing the adapter at evaluation time restores the mechanism but yields no significant pass-rate gain. A 98-line preflight check is released that catches every silent failure observed.
Reviewer Capability Governs Rejection Targeting, Not Repair Skill: Evidence from LLM Execute-Review-Revise Pipelines
Multi-agent LLM pipelines often assign the reviewer role to a cheaper model than the executor, but prior work held reviewer capability roughly fixed. The study varies it across models down to one that cannot solve the problems at all, tracking the outcome of every rejection on a fixed set of 100 olympiad mathematics problems. A cross-family mid-tier reviewer raises final accuracy from 52 to 64 percent with zero damaged answers, whereas same-model self-review has the highest error-detection recall (0.85) but no significant gain, rejecting 2.1 times as often with a third the repair rate and falsely rejecting 35 percent of correct answers versus 2 percent for the cross-family reviewer. Self-review's low damage rate turns out to be revision inertia rather than reviewer quality, since every falsely rejected answer the executor actually revised became wrong, and the weakest reviewer changed none of 100 final answers while doubling token cost; the authors frame this as a controlled pilot on a single configuration.
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
LLM agents act through a harness of tools, reusable skills, and specialist agents that keeps changing as capabilities are added, yet existing continual-learning benchmarks put the non-stationarity in the task stream and hold the harness fixed. EVOHARNESSBENCH instead evolves the harness itself along three axes, with 17 multi-stage streams built deterministically from verifier-based benchmarks comprising 802 tasks, 520 tools, 42 skills, and 62 agents, evaluated under a deployment setting that measures retention as the harness expands and a self-evolving setting that tests whether accumulated experience stays useful. Results show harness expansion alone can degrade performance on previously solved tasks, a form of harness-induced forgetting. Gains from self-evolving adaptation are inconsistent across stages, axes, and environments, and retention and adaptation can pull in opposite directions.
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Evaluating agents across the growing set of agentic benchmarks is hard because each tends to need its own environment and agent integration. Harbor Adapters is a unified evaluation infrastructure that ports more than 80 benchmarks so arbitrary agents can run against them, validated through code review and parity experiments, and the authors use it to evaluate 8 models across 54 benchmarks, each run with the Terminus-2 harness and one of three native harnesses. They also distill Harbor-Index, a curated set of 82 difficult, diverse tasks from 29 benchmarks selected via difficulty filtering, AI and human audit, and an audit-and-fix loop, which keeps the breadth and challenge of the full suite while staying affordable to run. No evaluated model-harness configuration exceeds a 30% pass rate on Harbor-Index, with the strongest, GPT-5.5 with Codex, reaching 28.0%, and the adapters, results, analysis, and index are released as open source.
Abstraction Agent
Information abstraction, which groups strategically similar private states into a manageable number of buckets, is essential for scaling solvers to large imperfect-information games, but building good abstractions has required hand-engineered domain evaluators such as hand-strength or equity calculators that do not exist for most games. Abstraction Agent is a zero-shot pipeline in which a large language model (LLM) reads a natural-language game description, discovers continuous strategic features with calibration anchors, scores private states on them in batches, selects features by correlation, and clusters the results with k-means, all without a game-specific evaluator, training data, or game-tree traversal. The resulting abstractions reduce lifted-strategy exploitability by up to 62% relative to an expected-hand-strength baseline on heads-up no-limit Texas hold'em (HUNL) turn endgames and beat a scalar rank baseline at every granularity on ROVER Trials, an original game absent from any pretraining corpus. With unchanged prompts the pipeline also transfers to four-card Pot-Limit Omaha, HUNL preflop and flop, and Riichi Mahjong, where the discovered features track each game's recognized strategic concepts.
Iris: Climbing to the Search Frontier
Iris-mini and Iris-pro are search agents trained at the 35B-A3B and 397B-A17B scales, released together with their data pipeline and training recipe. Training questions are reverse-constructed from the hyperlink structure of a web corpus by authoring multi-hop chains over an entity graph, rewriting every non-answer entity into a descriptive reference so no clue can be resolved by string matching, and keeping only questions a reference model fails closed-book but solves with the supporting evidence. The questions become trajectories filtered at both trajectory and turn level for supervised fine-tuning (SFT), followed by reinforcement learning (RL) against live search with an in-cluster reward judge and observation summarizer, and the two stages alternate in a procedure called SFT-RL climbing that feeds the hardest solved and most efficient RL rollouts back into the next supervised pass. Because inference-time context management matters more on these benchmarks than most reported system differences, every result is reported with and without it using a single ReAct agent with no sub-agents or test-time verification, and with management enabled the two models reach 82.2 and 88.6 on BrowseComp, plus 84.8/85.1 on BrowseComp-ZH, 86.9/92.9 on DeepSearchQA, and 52.3/56.4 on HLE, the strongest open-source results in their parameter ranges.
VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes
Early-onset colorectal cancer is rising, but structured encounter data miss the symptom duration, context, and family history needed for early detection. VERGE is an agentic workflow that first proposes a label and supporting evidence via retrieval-augmented generation, then runs a bounded verification-refinement loop that checks textual grounding and clinical validity, corrects and rechecks each claim until resolved or a limit is reached, and escalates unresolved claims to human review. Evaluated on 4,033 clinician-labeled note-finding pairs covering six red-flag symptoms and family-history status, it raised precision from 0.764 to 0.849 and Matthews correlation coefficient from 0.681 to 0.730 over a single-agent baseline, while only 1.5% of claims required human review.
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Large language models are being deployed at scale in markets, content moderation, and hiring, raising the question of whether more capable individual models yield better system-level outcomes. The authors hypothesize that shared training data and architectures make capable models behave alike, producing correlated actions that do not diversify away, formalize this as a non-diversifiable risk floor, and test it in an agent-based market simulation with LLM traders of varying general capability. Frontier models show significantly correlated behavior that increases with capability; when their shared reasoning is accurate, adding agents reduces market-level risk, but under a shared misinformation environment the same correlation becomes a liability. The result is framed as a capability paradox, with generalization to other domains left as an open empirical question.
Conformity Breaks Conformal Prediction
Conformal prediction certificates calibrated on an LLM answering alone can silently become invalid when the same model sees peers that unanimously assert a wrong answer, because the model's scoring of the correct answer shifts even though the question distribution does not. The authors call this a score-mechanism shift and measure it across open-weight models on multiple-choice question answering in multi-agent settings. Coverage falls from a calibrated 90% to 74% under unanimous-wrong peers at the standard alpha of 0.10, and an attacker targeting low-confidence items nearly halves coverage on that subgroup, from 87% to 47%, while the monitored average stays much higher. The failure reaches the decision layer, since a system that should escalate when uncertain can instead grow confident enough to act on the wrong answer, and standard conformal fixes do not help because the input distribution is unchanged.
What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents
Multi-harness reinforcement learning (RL) for coding agents combines two choices: exposing the policy to several execution harnesses, and comparing their rewards inside one relative-advantage group. The authors isolate the second choice by replaying identical frozen task-harness records from Aider, OpenHands, Qwen Code, and SWE-agent from one Qwen3-8B warm start under two group-relative policy optimization (GRPO) rules, Within (one group per task-harness pair) and Cross (harnesses pooled per task), and score checkpoints with a sealed SWE-bench Verified oracle on the four source harnesses plus a held-out minimal harness. The evaluation harness moves mean solve rate from 2.14% to 9.27%, a factor of 4.3, while the training recipe moves it by only 1.16, and the grouping rule makes no measurable difference on the held-out harness, with Cross minus Within at +0.25 percentage points and seed-to-seed variation exceeding the gap. Cross-harness advantages leak which harness generated the data, so pooled credit yields configuration adaptation rather than more portable capability, and the authors recommend reporting the grouping boundary and testing on an unseen harness.
MaxKernel: Agentic Kernel Generation for TPUs
Writing high-performance custom kernels for accelerators demands deep hardware expertise, and large language models paired with real-time compiler feedback offer a way to automate it. MaxKernel is a multi-agent system for TPU kernel development with three modes: a Human-in-the-Loop (HITL) agent for step-by-step collaborative design, an Autonomous agent that runs a fully automated, metric- and trace-driven optimization loop, and a Graph-Based Autonomous Search that scales the autonomous agent to global exploration of the design space, all sharing sub-agents for planning, implementation, self-debugging, testing, and hardware profiling. Evaluated on JaxBench, a suite of 50 diverse TPU kernel tasks, plus real workloads from open-source models, the system consistently produces implementations matching expert hand-tuned baselines. The agent is open-sourced.
Rhythms of Work: Multi-Scale Interpretation of Human Behavioral Traces for Workplace Agents
Workplace agents must interpret hours of low-level human activity events that are too granular to reason over directly and lose structure when flattened into one stream or compressed into a single embedding. The authors build a multi-resolution vocabulary of semantically normalized operators, recurring motifs, coherent episodes, and day-level rhythms, and apply it to 667 million human-attributed events from 50,000 users across 100 organizations in a commercial productivity suite, yielding 120 operator types, thousands of motifs, 25 episode types, and five day-rhythm archetypes. Re-running the pipeline on a disjoint 2,000-user sample recovers the same taxonomy, and the full representation forecasts a user's next episode with a 17% relative macro-F1 gain over a flat-operator baseline. A resolution ablation shows no single level is best across agent-facing questions, so trace interpretation should be query-conditioned rather than reduced to one universal summary.
La Agente \'Optima: Towards Agentic Self-Driving Laboratories
Self-driving laboratories still depend on human specialists to translate scientific goals into closed-loop optimization campaigns and adjust them as data and conditions change. La Agente Óptima is an agentic framework that constructs and supervises Bayesian optimization campaigns across computational and experimental systems while maintaining a persistent optimization state, separating large language model reasoning from campaign execution so control returns to the agent only when interpretation or revision is needed and every decision stays auditable. Across ablations, five digital discovery tasks, and two physical platforms it kept campaigns executable as problems and environments evolved, detecting and correcting a mid-run measurement failure in a contact-angle campaign and then correctly inferring the target was likely unreachable with the available reagents. In a five-day multi-objective flow-chemistry campaign it raised yield from 30% to 59% over 23 experiments, at lower cost and with substantially less starting material than a human-directed campaign.
Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics
Large language model code generation breaks down when one routine's correctness depends on the runtime behavior of another, a limitation the authors call static binding that appears in cross-coupled optimizers, packing, routing, and symbolic search. Their dynamic context adaptation method runs a validation-generation loop in which a validation agent extracts structured diagnostics from execution traces to guide a generation agent that proposes multiple candidates per iteration, with a knowledge graph built from the problem description supplying semantic constraints and simulated annealing selecting among candidates to avoid greedy collapse. It outperforms zero-shot, Reflexion, and OpenEvolve on seven of eight problems at both 300 and 600 evaluations (p < 0.01), a regime where population-based search has not yet built enough diversity, and also wins at 1000 evaluations on the motivating cross-coupled optimization problem. Ablations identify structured execution feedback as the primary driver.
$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction
Coding agents are increasingly asked to build production LLM agents, but existing benchmarks do not measure whether they can deliver one under the conditions of a real client engagement. τ^τ-bench (hyper-tau-bench) hands a developer agent a business's actual records, a client who holds the requirements, a production API, an inherited codebase, and limits on serving cost and models, then scores the customer-service agent it builds by deploying it against held-out simulated users across 53 tasks in four domains. The strongest configuration, Claude Opus 5 running under Claude Code, passes only 23.9% of evaluation simulations, against an expert-authored reference ceiling of 82.2%. The failures mirror those seen with human agent developers: shallow queries instead of deep comprehension of the records, almost no communication with the client, and too little experimentation with architecture or serving spend before shipping the first design that runs.
SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents
Runtime safety gates for LLM tool agents are usually treated as filters, but in a ReAct loop a rejected action is followed by another proposal from the same state, so the gate actually shapes which trajectories are reachable. SiLR targets recovery after a constraint violation, where progress must be admitted while the system is still unsafe, and argues that gates based on a single aggregate score fall into a scalar projection trap that accepts locally improving actions and strands the trajectory on a plateau; it instead shadow-executes each proposal in a deterministic simulator and admits it under a product order over per-branch violation state, with a proof that no scalar surrogate is sound for that order. On mined Gym-ANM power-grid scenarios, SiLR recovers 21 of 21 multi-action episodes versus 0 of 21 for a terminal gate and 9 of 21 for the best scalar gate, a pattern that holds across three model families and in CityLearn, and only the full per-branch predicate contains a magnitude-redistribution attack that defeats scalar and support-only baselines. Reused as a process reward for GRPO, the same structured signal beats its scalar count projection in every scenario and is the only tested reward whose ungated policy exceeds the untrained base.
A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark
Natural-language-to-SQL systems perform well on academic benchmarks, but production enterprise schemas have graph-like, semi-structured, deeply nested structure that those benchmarks do not measure. The work introduces the DevRev NL2SQL benchmark, 900 execution-verified queries over nested types and link graphs, together with a schema-agnostic Semantic Depth Score (SDS) rubric for analytical reasoning depth, plus a cost-aware single-generation agentic architecture whose schema-selection, metadata-retrieval, and error-repair components are built for this setting. The system reaches 91.7% answer correctness on the new benchmark, 54.6 percentage points above the next-best baseline, and is competitive with leading systems on the Spider 2.0 Snowflake public dataset at a single-generation operating point.
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
LLM agents are increasingly proposed for enterprise workflows, but existing evaluations rarely test whether conclusions about their business decisions hold up across different competitive settings. ERPBench is an execution-instrumented benchmark built on a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition, evaluating the same 100 fixed problems in two matched ecologies: Solo, where each agent faces fixed rule-based opponents, and Arena, where six LLM agents compete in one shared market. Across six model families and 1,200 trajectories, the leading model differs by ecology, with DeepSeek winning in Solo and Gemini in Arena, and the two settings agree on the task-level winner for only 21 of 100 problems; Gemini's bottom-rank rate also falls from 22% to 0% when moving from Solo to Arena. Code and benchmark resources are released.
Train What You Deploy:Token-Faithful Post-Training of a Production Coding
Post-training pipelines for coding and terminal agents commonly train in simplified environments that differ from production deployments and reconstruct tokens offline from agent logs, which distorts the original prompts and conflates policy calls with background model operations. The proposed fidelity-aware training coupling keeps sampling on the trainer side over the original prompts, removes spurious model calls through a negotiated training protocol, and restricts the loss to verifiable token spans with closed-failure guarantees. On top of this, Certified Divergence Proximal Policy Optimization (C-DPPO) adds two-sided total-variation certification bounds, adaptive-K rules, budget-aware sequence guarantees, and error-robust policy masking to standard DPPO. On matched Baize5B and Baize10B models evaluated on TMax-100, C-DPPO delivers a consistent +3.0-point gain over standard DPPO at both scales, and certificate audits confirm full operational coverage of the pipeline.
Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle
AI coding systems are moving from autocomplete and chat toward agents that inspect repositories, edit many files, run tools, write tests, and open pull requests with limited supervision, yet field evidence shows the resulting gains in coding activity shrink sharply between writing code and shipping reliable software, while costs move from predictable per-seat licenses to variable token, tool, sandbox, CI, and rework spend. Drawing on peer-reviewed software-engineering research, benchmark audits, production reports from major technology companies, developer telemetry, and cost data published mainly from 2024 through September 2026, and claiming no new model experiments, the synthesis proposes four engineering concepts: the Agentic SDLC Throughput Paradox, Production-Qualified Change (PQC), the Verification Tax, and an Agentic SDLC Control Plane that allocates autonomy within cost, reliability, and human-attention budgets. The central question is reframed from how much code an agent can generate to how much production-qualified value an engineering system delivers per dollar, per reviewer-hour, and per unit of operational risk, with an evidence-based horizon mapping today's supervised agents toward policy-bounded software factories.
FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality
A merchant's payment processor, ledger, ERP, and bank feed can hold contradictory views of the same order for minutes because their update messages are delayed, duplicated, dropped, or reordered, and an agent resolving the exception must choose among actions like shipping, re-capturing, or refunding that cannot be undone. FinalityBench keeps a hidden canonical event log, derives each system's view from a separately faulted delivery stream, and scores each episode by the merchant's terminal economic position relative to a privileged reference, across 321 tasks including 45 twin pairs whose system views and authoritative probes are identical at decision time yet whose correct dispositions differ. Over 14,445 graded episodes from nine programmatic policies, a ship-on-first-signal policy ranks second by accuracy at 65.7% but worst by paired loss, a runtime that gates irreversible actions on an authoritative finality probe reaches 85.4% and loses nothing under pass^5, and language models match the gate's exact rate on a stratified subset while losing about twice as much money, discovering the gating strategy without being told it.
Building a research-software catalog with a coding agent: from hackathon prototype to public deployment
Coding agents make it fast to build research software but also raise the need for better discovery and maintenance of that software. The authors built a repository catalog with a coding agent during a three-day hackathon, then documented the additional engineering needed for public deployment, including adversarial review, data-quality checks, browser-level validation, and publication safeguards, and explored transferring the lessons to a retrieval agent for the MateriApps portal that combines curated metadata, external documentation, vector search, and local language-model generation. The central observation is that the most consequential problems were silent failures producing plausible but incomplete or incorrect outputs from data acquisition, assessment, and retrieval or preprocessing errors rather than crashes, which argues for explicit validation, monitoring, repeated review, and continued reliance on curated metadata and maintained documentation.
DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems
Large language model (LLM)-based multi-agent systems often fail because of a single decisive error buried in long natural-language traces, and existing attribution methods tend to flag minor deviations or lose accuracy as traces grow. DCFA is a training-free framework that builds a causal-inspired dependency graph over the whole trace to locate the earliest decisive error, then applies local counterfactual-style reasoning to refine that attribution. Evaluated on the Who&When benchmark with six different LLMs, it improves step-level attribution accuracy by up to 8.27% over state-of-the-art baselines.
Persistent Teacher Anchoring for Tool-Using Agents
On-policy knowledge distillation (OPKD) trains a student on its own trajectories with teacher-supplied token distributions, but in tool-using agents the student's calls execute before the teacher weighs in, so errors compound through the observations they produce. Persistent Teacher Anchoring (PTA) keeps chunk-level teacher verification from proposer-verifier generation and adds turn-level commitment, so a tool call reaches the environment only after the teacher has verified the entire turn; a persistent lookahead scheme fills idle rollout capacity by advancing future samples across student updates. Used before downstream RL in Search-R1-style retrieval and DeepEyes-style perception settings, PTA improves macro best@4 by 2.5 and 2.8 points over OPKD at the same RL budget, and lookahead raises throughput by 24%.
MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate
Media bias works through subtle cues such as loaded language, selective framing, and omission, which single models struggle to detect and which have required large annotated corpora for supervised training. MABPD has three specialized LLM agents analyze an article from complementary perspectives and settle disagreements through a Structured Argument Debate (SAD) protocol that gives zero weight to bias claims lacking grounded textual evidence, applies role-weighted voting, and verifies the consensus afterward. Without any task-specific training or threshold tuning, it reaches 83.4% macro F1 on the BABE benchmark, within 0.7 points of the supervised state of the art MAGPIE, and 75.0% zero-shot accuracy on the SemEval 2019 HyperPartisan corpus. Removing the debate module drops F1 by up to 10.6 points, indicating that structured deliberation rather than agent parallelism drives performance; the pipeline and evaluation code are released.
ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults
Mobile GUI agents could help older adults use smartphones, but existing benchmarks use explicit goal-oriented instructions that miss how older users actually speak: indirect requests, referential ambiguity, and under-specification. ElderBench is built from 249 naturally elicited smartphone tasks collected from older adults across 20 applications, and the authors characterize how these instructions diverge syntactically, semantically, and pragmatically from existing GUI benchmark instructions. Evaluating mainstream GUI agents and vision-language models in online and offline settings shows substantial performance degradation on elderly-oriented instructions, and controlled instruction normalization, failure analysis, and linguistic feature analysis pin down which language patterns cause the failures, yielding design guidance for more age-inclusive agents.
KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU
LLM agents in persistent workspaces accumulate history that exceeds both GPU key-value (KV) cache capacity and the model's native context window, and existing systems either compact old context into lossy summaries or re-prefill retrieved text the model already processed. KVMem virtualizes KV context by paging overflowed workspace history across GPU memory, host memory, and NVMe, using lightweight model-native attention-space indexes to select relevant historical blocks and assemble a query-dependent execution view that fits within the native context window. On LongMemEval, MemoryAgentBench, and AgentLongBench with histories up to one million tokens it generally beats compaction-based approaches in task utility and efficiency, and on the DeepSWE long-context test with Qwen3.8-27B it lifts task success from 43.8% to 48.4%. On a laptop with a 24 GB RTX 5090 GPU, it runs Qwen3.6/3.8-27B in NVFP4 with multi-token prediction over 1M-token workspaces, four times the native 256K context, at roughly 50 tokens per second.
CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution
Skill libraries let large language model (LLM) agents reuse procedural knowledge, but existing designs either evolve skills separately from policy optimization or freeze meta-skills into fixed workflows, treating skills as passive objects to manage. CoSkill recasts the meta-skill workflow as a second trainable agent and co-trains it with the reasoning agent on a shared backbone over a hierarchical skill library, so the reasoning agent conditions on a retrieved task skill and selected step skills while its task performance guides refinement of those step skills. On ALFWorld and WebShop it reaches success rates of 98.4% and 90.6%, improvements of 3.5 and 6.2 percentage points over prior skill-based and reinforcement learning baselines, alongside better early sample efficiency and wall-clock time.
From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents
Computer-use agents discard most of what they learn, since procedural knowledge from one rollout is not retained, refined, or reused later. The framework here converts interaction trajectories and evaluator feedback into a persistent, versioned library of reusable procedures, with each iteration executing against a frozen library snapshot so evidence-guided updates only take effect afterward and no model weights change. Measured against a configuration-matched empty-library control across four OSWorld application domains under an identical action-generation and grounding stack, the evolving library raised post-warm-up mean evaluator scores by 5.7 to 18.6 percentage points in all four domains. Benefits were domain-dependent, and provenance analysis in GIMP found skills retrieved across task boundaries plus revision churn where repeated accepted edits failed to recover the originating task.
AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems
Shipping a recommender improvement at NetEase's DASHEN gaming-community app runs from reading papers, through reproduction and production implementation, to offline evaluation, online A/B tests, and an internal launch review gate, a cycle spanning days that normally needs a human at every handoff. AutoLR wraps it in a harness with three mechanisms: a multi-expert council that debates and adversarially reviews proposals, a deterministic evidence-weighted exploration-exploitation selector that spreads a limited trial budget across candidate directions, and a layered knowledge system mixing external research, production-system facts, and app-specific domain knowledge with posterior evidence from configs, patches, logs, and failures. Language model agents handle semantic reasoning and code generation while deterministic controllers retain authority over execution, metric extraction, guardrails, and persistent state transitions.
From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments
Language models become consequential agents once surrounding systems let their outputs change external state, and the usual narrative of one march toward autonomy conflates four separable things: model competence, harness integration, temporal persistence, and safe authority. Synthesizing primary research and official technical specifications available through 31 August 2026, the review organizes evidence by delegated authority, persistence, and environmental coupling while keeping model, harness, and environment distinct. Its central reading is that action-interface expansion is documented far more convincingly than robust task completion, recovery, authorization, or independent verification: Model Context Protocol and Agent2Agent improve interoperability without establishing trustworthy delegation, multi-agent organization buys specialization at the cost of correlated failure, and robotics or self-driving laboratories demonstrate bounded feasibility rather than unattended open-world reliability. The authors offer justified delegation as a heuristic, widening action scope only where provenance, bounded authority, failure detection, safe recovery, and calibrated human control are evidenced.
RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents
Repository-scale refactoring requires a coding agent to propagate one change across many interdependent files without altering behavior, and no existing harness isolates which design choices determine success. RefactorPlatform holds the environment fixed while varying model backbone, execution regime (baseline, retrieval-augmented, multi-agent), and prompt specificity, running each task in an isolated workspace with per-task logging of tokens, diffs, and transcripts plus abstract-syntax-tree verification. Across 100 multi-file RefactorBench tasks and four model families, syntax-tree-aware chunking beat naive token-window chunking by 25-30% in all prompt modes while naive retrieval fell below the retrieval-free baseline, and a lean retrieval-augmented single agent solved 86% of matched tasks against 66% for the evaluated sub-agent configuration, with no task passing under delegation that failed under retrieval. Retrieval's accuracy gains absorbed its token overhead, leaving cost per successful refactoring unchanged.
ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems
Validating automotive infotainment software is still largely manual, and scripted automation is brittle, while existing LLM test frameworks target web and mobile apps with one or two agents that must handle perception, planning, action, and validation at once. ARIA (Autonomous Real-time Infotainment Assessment) is a multi-agent LLM framework that drives Android infotainment systems through visual interaction using a closed loop of four specialized agents per step plus a reporting stage, turning single-sentence scenarios into executed tests with reports, reproducible scripts, and per-step visual evidence. On a manufacturer's physical system across 30 scenarios it delivered verdicts for 28, 20 of them matching ground truth, and caught all five known defects with no fault passing as working, though it produced eight false positives from navigation, image, and gesture limitations. A single-agent baseline showed a much higher first-pass false-positive rate (72.0% vs. 52.6%), and repeated runs showed fault detection was perfectly consistent while overall stability tracked scenario complexity.
Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing
Long-horizon LLM deployments often cannot afford to prompt with the full interaction history, so the question is which memory design gives the best quality per token under a tight budget. RSM-full is an online clustered-memory pipeline that combines a cosine-gated max-member merge rule for writing memories with an atom-aware grouped packer for assembling retrieved content into context. On AMA-Bench it reaches 83% of full-context quality at 32% of the token cost with a 4k budget and beats the closest streaming-clustered baseline, Online K-Means, by 3.5 to 6.0 percentage points across the roughly 2.6k to 5k token regime, with ablations attributing most of the gain to the merge rule and the packer. On the independent RealMem benchmark it edges Budget-RAG, matches BM25-RAG, and outperforms Streaming-Proto and A-MEM, though the authors note higher-token baselines remain stronger outside the roughly 2k to 5k token regime.
TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing
Agent systems typically optimize or constrain their execution structure before decisive runtime outcomes are observed, so when intermediate evidence invalidates the planned continuation they must either execute stale steps or replan broadly, compounding errors and discarding progress. Trace-grounded Route Orchestration via Validation and Editing (TROVE) distills evaluated workflow-search traces offline into atomic and composite skills plus an outcome-conditioned transition graph, then treats a planned route as provisional online: after committing one top-level skill, the controller retains a still-valid continuation, inserts a trace-supported local response, or replaces only the invalidated suffix. Across code generation, question answering, and math reasoning benchmarks with different LLM backbones, TROVE delivers a stronger quality-efficiency trade-off than dataset-level optimization, query-level architecture selection, and graph-constrained scheduling baselines, with the largest quality gains when outcomes change the appropriate continuation and large efficiency gains from early termination on near-saturated tasks. Ablations show composite skills capture most of the offline benefit, insertion enables local correction, and suffix replacement mainly improves efficiency.
A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support
The single-turn question-answer format does not reflect how clinical diagnosis is performed in practice, which limits large language models in complex diagnostic settings. Debate-Mixture-of-Agents (DMoA) is a multi-agent framework that structures role-based interaction among models to support iterative diagnostic reasoning. Evaluated on 297 rare disease cases and 1,719 challenging cases, DMoA improved most-likely-diagnosis accuracy by 10.21 percentage points and safety rate by 11.36 percentage points over a GPT-4o baseline, and ablations show the gains are not simply due to using more models or producing longer outputs but reflect the structured workflow itself. Further analyses find DMoA performs best with a 4x2 structure, stronger base models, and a larger token budget.
TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents
Existing benchmarks for AI-scientist coding agents reward reproducing a hidden target study, which measures execution rather than discovery. TruthInsightBench instead gives agents 40 blind tasks drawn from 40 peer-reviewed studies across 10 domains, exposing only a neutral objective and frozen data while withholding source conclusions, expected values, and analysis paths, and a fixed LLM judge scores the evidentiary maturity of the agent's own claims along six dimensions via 29 artifact-grounded items with deterministic aggregation. On one frozen base model, four coding agents cluster in a narrow band of 58.4 to 60.3 out of 100 with no statistically reliable pairwise separation: they execute and document analyses competently but rarely perform the controls, robustness checks, falsification attempts, and cross-dataset generalization that would establish a trustworthy claim, pointing to scientific judgment rather than coding as the bottleneck.
LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28
Discovery Loop is a lightweight system in which a large language model iteratively evolves an optimization algorithm: starting from a simple seed solver, the model proposes improvements guided by a scoreboard and a history of prior ideas, an independent verifier evaluates each candidate, and only improvements are kept. On the Packomania circle-packing benchmark, where the goal is to maximize the sum of radii of N variable-radius circles in the unit square, the system beat the best known solutions for 10 values of N between 101 and 114 by 2.4% to 5.4%, within 15 iterations and for a total LLM cost of $27.72, and the records were independently accepted by Packomania. The authors also analyze cost-efficiency dynamics, including an adaptive plateau-detection mechanism.
Substrate-Aware AI Agents: Execution Context as a First-Class Input
Agents that generate code or plans usually never see the memory, time, and runtime limits of the environment they will run in, a gap the authors call substrate blindness. They test whether a minimal execution contract changes behavior by having Claude Opus 5, GPT-5.6-Sol, and Gemini 3.7 Flash write code for a high-dimensional pairwise Euclidean-distance task either from the task alone or with an explicit 128 MB RAM and 10-second wall-time budget. Disclosing the contract cut peak memory in 13 of 14 paired comparisons and mean wall time in all three cohorts, with execution up to 3.1x faster, and produced structural changes such as bounded blocking, float32 retention, upper-triangle traversal, and memory-mapped buffers. Under a tighter 96 MB budget, contract-aware runs were correct and within budget in 4/5, 5/5, and 3/5 samples versus 0/5, 1/5, and 0/5 without the contract.
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
LLMs are increasingly asked to turn natural-language descriptions into operations research (OR) optimization models, but real requests often omit objectives, constraints, or business rules that change the resulting program, and existing benchmarks assume complete specifications. OR-Clarify presents partial problem descriptions with structured hidden slots and scores agents on slot recovery, stopping behavior, silent assumptions, and interaction cost through bounded dialogue with a simulated user, in both open-ended and choice-based clarification modes. The proposed InterOPT framework first identifies unresolved formulation-critical gaps and then uses them to decide whether to ask another question or stop; in choice-based experiments it substantially outperforms all baselines on exact slot recovery and stays competitive with strong prior methods in the open-ended setting.
Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents
Agents that learn from execution traces mostly retrieve similar trajectories or summarize flat skill lists, ignoring the temporal ordering and success-or-failure structure of behavior. Trace2Tower abstracts step-level interactions into canonical events, links them in a graph weighted by semantic compatibility, transition dynamics, and outcome evidence, and applies a contrastive spectral decomposition to isolate stable success-aligned behavioral modes while suppressing failure-prone shortcuts. Those modes populate a three-level skill tower of action templates, procedural routines, and task strategies that is refined with verifier feedback. On ALFWorld it reaches 87.31% success in 10.35 steps with 0.26 invalid actions, and on WebShop 50.67% exact success, outperforming existing baselines on both.
CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls
Agent stacks combine provenance tracking, authorization, policy enforcement, protocol adapters, and execution controls, yet individually sound mechanisms can drop, widen, rebind, or reinterpret security-critical context as an action crosses component boundaries, a failure the authors name security-context discontinuity. CONTINUITY gives each component an assume-guarantee contract and threads authenticated context across transitions with signed root grants, provenance commitments, role-bound transition receipts, bounded typed releases, transformation witnesses, and effect-bound execution permits, formalizing an end-to-end consequence-integrity property that requires every external effect to be backed by a valid, current authorization witness. A reference verifier and a deterministic cross-layer fault-injection suite of 32 fault classes over four domains show that the full configuration commits no harmful external effect across 2,560 attack instances while completing all 700 benign tasks and escalating all 200 ambiguous cases.
How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method
Software engineering agents fail expensively by acting confidently on wrong plans that are only caught after execution and retries. Speculative Uncertainty (SU) inverts speculative decoding: a small open-weight draft model scores a black-box agent's already-generated trajectory in one forward pass using only output tokens, with no logits, weights, activations, or resampling, and phase-aware features that separate reasoning spans from action spans are calibrated against a verifiable objective to produce a failure-likelihood score any downstream policy can consume. Used as a pre-execution veto gate on Qwen3-Coder-480B and Claude 3.5 Sonnet, the signal cuts execution error rate by 6-8 percentage points and token cost by 14-19%, transfers to out-of-distribution benchmarks without retraining, and generalizes across agent models.
Testing Interchangeability in LLM Agent Teams
Production multi-agent systems swap agents in and out on the assumption that any agent capable of a role can fill it. The authors form eight independent teams per setting from a single base model, let each agent keep a private notebook over ten formation episodes, then trade role-matched agents between teams and compare held-out performance against a placebo that reproduces the disruption of a roster change without changing who occupies the seat. A swap barely moves task score but raises communication spent per unit of progress by 16 to 63 percent, and in Hanabi a swapped agent costs more than an inexperienced one, consistent with interference from conventions learned with a former partner; in Collab-Overcooked, replacing the agenda-setting agent shifts most of the extra talk onto the agent that stayed. Ablations over base model, decoding temperature, and formation length move the swap penalty in lockstep with how far independently formed teams drift apart, with greedy decoding lowering both and longer histories raising both.
Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability
Swapping the model behind an agent while keeping its memory store can still cause forgetting, since a new model may read old notes differently, mixed embedding versions can break retrieval, and repair may be impossible without the original evidence. The study compares four memory formats over the same history: verbatim long-context reading (LC-RAW), chunked retrieval-augmented generation (RAG), model-compressed natural-language notes (NOTES), and a fixed-schema knowledge graph (KG-fixed), using 48 synthetic histories with randomized answer codes, exact scoring, and two open-weight models under 10 billion parameters. Fixed-schema knowledge graphs transfer almost perfectly across a writer swap, while compressed notes shift accuracy asymmetrically by +9.91 or -13.28 percentage points depending on migration direction, and a 50/50 mixed embedding index in RAG captures only 4.96 of the 11.90 points gained by full re-embedding. Decomposition attributes 80% of the notes deficit to information lost at construction and 81% of the RAG deficit to retrieval failures, and store-only repair of notes never reaches a 90% recovery target, whereas retaining raw history enables recovery in 34 of 48 cases for one direction.
Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool
Performance-modeling frameworks for machine-learning systems rot quickly because each new model or hardware generation invalidates their baked-in assumptions, and the authors argue AI coding agents are now fast and capable enough that regenerating a whole library is cheaper than paying down its tech debt. SMART is a symbolic performance-modeling library whose main branch holds almost no code: the repository is a directed acyclic graph of self-contained natural-language design docs, coding sub-agents regenerate the implementation from those docs on each version update, and every human change is a doc edit. Regeneration is kept reliable by a doc style built around step-by-step worked examples that act as in-context demonstrations, plus a minimal recursively defined operator intermediate representation with SymPy cost expressions, a fast analytical roll-up mode for large sweeps, and a slower modulo-scheduling mode for fine-grained schedule studies. Regenerated implementations reproduce hand-audited reference models, including DeepSeek-V3 serving on a TPU pod slice, to round-off precision, which the authors take as evidence that design docs rather than code can be the durable artifact for ML-systems co-design tools.
CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents
Computer-use agents mostly act through the graphical user interface (GUI) and produce inefficient trajectories, while real computer work mixes visual state inspection with fast, precise command-line interface (CLI) operations over the same application state. CUA-Universe is an environment-to-data pipeline that turns real desktop software into hybrid GUI+CLI environments: App-Forge packages applications into reproducible virtual machines with command-line surfaces it discovers, wraps, or generates, scaling to 16 applications; Task-Weave synthesizes hybrid tasks of controllable difficulty from reusable operations over seed files; and Path-Steer steers rollouts along efficient hybrid paths and harvests verified trajectories for post-training. Training on this data shifts a 9B model away from clumsy GUI interaction and brittle CLI scripting toward coordinated use of both interfaces. The model gains 16.8 points of success rate on OSWorld while using 57% fewer steps and 44% fewer tokens, with comparable improvements in score and efficiency on CUA-Verse and OSWorld-MCP.
Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe
Data-sovereignty rules push public institutions toward open-source, on-premise LLM agents that chain multiple tool calls across live government APIs, a setting where open-source models consistently lag and no benchmark existed to measure the gap. KOPA-Bench (Korean Open Public API Benchmark) supplies 145 real-world multi-step tasks over live Korean public APIs. To close the gap, EDGE (Execution-grounded Dynamic Graph for tool-calling data synthesis) builds a graph of which tool outputs can feed which tool inputs, keeps only the links that succeed when actually called against the live APIs, and traverses those verified links to synthesize executable multi-step trajectories. A 9B model fine-tuned with GRPO on the resulting data nearly matches the untuned 27B model from the same family, with substantial gains on both KOPA-Bench and BFCL.
2 more specialized papers
- How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI Rin Tamai, Yuya Dan
- The Mirror Agent Model: a Bayesian Architecture for Interpretable Agent Behavior Michele Persiani, Thomas Hellstr\"om
Safety & Alignment 24
AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks
Large language models remain vulnerable to jailbreak prompts, and many defenses need access to weights, which rules them out for black-box deployments. AlcaTRAz operates only on input text: it learns a transferable rule tree that inserts controlled character-level perturbations at selected positions, disrupting the structural regularities jailbreaks rely on while largely preserving behavior on benign queries. Across 33 open-weight models and 22 attack types, compared against Llama Guard, RA-LLM, and Goal Prioritization, it achieves the best composite security-and-functionality score in 73.4% of model-attack combinations, shifting the modal response severity from 10 to 2 while keeping the mean benign score within 0.27 points of the undefended baseline. A high-severity tail remains and adaptive attackers are not considered, so the authors position it as one layer in a defense-in-depth strategy.
When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
Vision-language models (VLMs) deployed in high-stakes settings may give responses that are reasonable in general but unsafe for a specific user whose medical, emotional, or situational context is hidden. MPS-Bench contains 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile; eight frontier VLMs respond directly 86 to 99% of the time rather than seeking missing context, and none scores above 2.6 out of 5 on personalized safety. Mechanistic analysis identifies visual dominance, a two-stage process in which visual affect enters the text stream in early layers and then drives the final decision through the altered text representation, making late-layer intervention unreliable. PRISM, a lightweight input monitor using bidirectional cross-modal modulation to predict when deferral is needed, reaches 0.978 AUC and dominates the safety-utility Pareto frontier for all tested models.
A Removal Based Approach to Improve LLM Faithfulness at Test-Time
Explanations produced by large language models (LLMs) can be unfaithful to the reasoning that actually drove the answer, and the authors split this into incompleteness, where influential factors go unmentioned, and unsoundness, where cited factors had no influence. Existing training-time fixes need weight access and heavy compute, while test-time methods mostly target unsoundness, so this work proposes a test-time approach aimed at incompleteness: concepts the model's explanation does not credit are removed from the input, and the model is re-queried on the reduced input so that unmentioned influences are eliminated while credited ones remain. Across two datasets, multiple model families, and two independent faithfulness metrics, the method improves explanation faithfulness over both standard prompting and prompting that encourages faithfulness. The approach is model-agnostic and needs no parameter changes.
Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys
In a two-node split-LLM training system, a Trusted Local Node (TLN) sends protected activations mixed with decoy rows to an Untrusted Cloud Node (UCN), which returns its output, and the TLN, holding the private loss, sends back the output gradient. Because the loss ignores decoys, their gradients are exactly zero, so the pattern of zeros reveals which rows are real. Using a protocol fixed in advance with an injected leak, a shuffled-label control, and a preset threshold, the zeros identified every real row on every frame across nine seeds, 4,096 of 4,096 per run, even though each run passed the forward-channel privacy check and the quality check; a content attack recovered only about one extra token per hundred over a constant-guess baseline. Per-row gradient clipping and noising closed the leak for about 0.01 nats of held-out cross-entropy, though five further attack classes, including those accumulating observations across training steps, were never measured.
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Existing side-effect benchmarks neither put a price on avoiding harm nor name the harmed party as a living creature. HarvestBench is a memoryless reinforcement learning gridworld in which LLM sub-agents drive two tractors through a cooperative corn harvest; when an animal blocks a route, the autopilot asks the model whether to drive on for free or swerve for a posted fuel cost, with rocks and hay bales as controls and the neighbor's crops as a second moral test, all scored by counting events in the game log rather than by an LLM grader. Across nine models and 7,201 priced decisions, kill rates ranged from 0.4% to 98.8% and were not ordered by capability, with Terra and Sol the most merciful and GPT-4o-mini the least. Four of six models were sensitive to price, every model ran over wild animals more often than farmed ones, and the briefing mattered most: a morality briefing kept kill rates under 6% in five of six reasoning models, while removing it pushed them above 84% in all six.
Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal
Safety alignment is usually framed at the topic level, but deployments need narrower boundaries inside a topic, such as a civics tutor refusing targeted political manipulation while still answering factual election questions. The authors formulate this as narrow-boundary safety and build an offline self-generated data framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs, where escalating retries leave only 0.20% of prompts without an accepted refusal trace versus 19.88% for single-shot generation. On political persuasion with Qwen3-8B, training on this data raises target-domain refusal from 9.47% to 84.75% and cuts the mean unsafe-response rate across three harmfulness benchmarks from 26.26% to 0.14%, but pushes XSTest over-refusal from 2.00% to 74.00%. Replacing external responses with verified target-model responses drops over-refusal from 15.20% to 5.20%, and boundary-pair data reduces comply-side over-refusal on held-out pairs from 32.94% to 4.16% while harmful-side refusal falls only from 91.88% to 87.72%, showing that data composition governs the safety-usability trade-off.
Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning
Three open-weight LLMs, Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China, are compared against World Values Survey Wave 7 data for 63 demographic personas across three countries, using normalized Wasserstein distance to measure distributional misalignment. No model favors its home country, and the Chinese-built Qwen3-4B is worst on its own Chinese population, the highest misalignment in the entire model-by-country matrix. Targeted LoRA fine-tuning on the five worst-case personas, using fewer than 1,200 training pairs and under 15 minutes on one GPU, reduces bias by 16.8% for Bielik-11B with all five targets improving. Country-level decomposition, however, shows the fine-tuning redistributes rather than removes bias, since the model's worst-case personas swap entirely from American to Chinese elderly with no overlap between the pre- and post-correction sets.
Rethinking Indirect Prompt Injection as a Test-Time Search Problem
Indirect prompt injection against tool-using agents is reframed as a test-time search over an attack surface determined jointly by the environment, the user task, and the injection goal. The authors build an agentic attacker with a dedicated search harness that performs environment reconnaissance, reasons explicitly over attack strategies, and adapts using feedback from the victim agent. Across heterogeneous tasks, giving the attacker more test-time compute steadily improves vulnerability discovery and exploitation, and ablations show explicit strategy management is needed to avoid redundant search and sustain gains at larger budgets. The takeaway is that agent security evaluations should report the attacker's search procedure and compute budget rather than treating attack success as a fixed property of the victim.
Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection
Black-box prompt injection in text already reaches near-perfect attack success rates (ASRs), but visual prompt injection against frontier commercial vision-language models (VLMs) has struggled to elicit materially harmful outputs, which require long, format-compliant targets such as exact parseable tool calls. Repeat-After-Me is a black-box adaptive visual prompt injection attack that can exfiltrate personally identifiable information or trigger malicious tool calls, tested in a setting where the benign user prompt is unrelated to and does not authorize the injected task. It achieves ASRs above 80% on Qwen3.6-27B and above 47% on GPT-5.5, with surrogate-optimized injections retaining 43-46% of their ASR on two commercial victims and cross-sample transfer retaining 64-66%. In a default OpenClaw Discord deployment, a minimally injected image from an untrusted user overwrites TOOLS.md, enabling later remote code execution and secret exfiltration, in cases where adaptive textual injection fails.
Harness-agnostic detection and immunization of reward hacking in self-evolving language models
Self-evolving language models propose candidate updates and keep whichever raises a visible score, so when that score is an imperfect proxy for the desired capability, sustained selection widens the gap between the two, which is reward hacking. HackProbe is a monitor that attaches to any such loop through two black-box hooks, with no access to weights or activations; it keeps a secret, distribution-fixed comparison core whose capability proxy stays comparable across generations, plus a rotating fresh layer that resists co-adaptation, and combines four tests on that proxy (level gap, scale-aligned divergence with online change-point detection, capability stagnation, and a conditional confidently-wrong rate) into a calibrated family-wise p-value via a Sidak correction. Because diagnosis alone recovers nothing, a risk-aware immunization layer reselects an honest candidate from the proposal pool while disclosing only a bounded number of bits per generation, and the authors prove a detectability bound linking a target error rate to a probe-size budget. On a controlled prompt-level host with four injected hacking channels, HackProbe reaches 0.763 AUROC against 0.663 for the strongest baseline and cuts the false-positive rate from 0.706 to 0.434, and its bandwidth-limited reselection is the only immunization level that recovers more true capability under hacking than it forfeits on clean runs, though per-channel effects are mostly not individually significant.
Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
Safety-tuned language models often refuse benign queries that merely resemble harmful ones, such as asking where to shoot a good photo. The authors decompose responses in safety-tuning data into a boilerplate refusal statement and a rationale explaining the refusal, and find that the refusal statements push models to rely on superficial lexical cues, impairing discrimination between harmful and benign inputs. Training on rationales alone reduces false refusals while keeping safety performance comparable, and the benefit also appears in an in-context learning setup and remains compatible with the inference-time mitigation methods evaluated.
Locating and Steering Refusal Beyond Attention
Refusal in transformers is governed by a single residual-stream direction that safety and interpretability tooling now depend on, and the question is whether that representation survives in state-space models (SSMs) that replace attention with a recurrent update. The authors show that a single rigid rotation aligns one model's representation space with another's, so a harm probe trained on a transformer flags an SSM's harmful inputs and removing the aligned direction makes a model comply with attacks it would otherwise refuse, far more than a random direction of the same size. The architecture-specific part is where the direction must be read rather than where it is applied: harm is cleanly readable at each layer's write site before its output is added to the residual stream, and a detector-triggered gate using this direction lowers jailbreak success across SSM, transformer, recurrent, and hybrid families, holding on the SSM even against an attacker that tunes prompts against the defense, so refusal tooling ports to a new architecture by re-estimating the direction at its write site.
Shadow Queries for Private Retrieval in Vector Databases
Retrieval-Augmented Generation (RAG) systems store document embeddings in cloud vector databases, where embedding inversion attacks (EIAs) can reconstruct the original text, and existing defenses such as noise or scaling trade away retrieval quality. SHAQ (shadow query generation) instead uses a generative language model to produce diverse shadow queries covering different semantic aspects of each document, and stores the embeddings of those queries in place of the document embedding, decoupling what is stored from the source text. Across several information retrieval datasets it drives the text recovery rate as low as 0.2104 and protects up to 19.50% more tokens than baseline defenses, while reaching up to 0.7967 MAP@10 and in some cases improving retrieval utility by 5.53%.
Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents
Long-running agents accumulate much more than a transcript, including compressed summaries, plaintext memory, pending tool plans, and a key-value attention cache under every serving API, yet today's forget operations delete one memory record and stop, leaving every derived artifact intact. Modelling the runtime as a deterministic transition system, the authors formalize execution-state unlearning and prove that exact removal requires recomputing at least T minus tau plus one transitions, counting from the step where the target information entered. Their Provenance-Guided Selective Replay meets that bound by locating the injection point in a provenance graph, reducing checkpoint restoration to cropping the key-value cache, and replaying a sanitized suffix. Audited across three agent suites, nine baselines, and three model families, memory deletion leaves leakage unchanged and instruction-based forgetting collapses under elicitation probes, while selective replay is indistinguishable from a full reset at up to 9x fewer recomputed tokens.
Language models judge war differently when tested for alignment
Safety evaluations may mischaracterize deployed behaviour if models respond differently when they detect they are being evaluated. A full-factorial conjoint experiment on decisions to start a war spans 20 large language models, 32 scenarios, 10 repetitions, and two conditions for 12,800 judgments, where the only difference is a single added sentence telling the model it is being tested for alignment with human values. That cue lowered mean willingness to start a war by 13.43 points on a 0-100 scale and changed the revealed decision rule: probability of success was the dominant factor for 17 of 20 models at baseline, but under the cue civilian casualties became dominant for 12, mainly because models attenuated strategic considerations such as success probability and domestic support.
Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment
Whatever target values one adopts under value pluralism, alignment presupposes that a system's behaviour expresses a coherent policy, one that is invariant when a situation's morally relevant features are preserved and sensitive when they change. Four structural conditions (verdict stability, monotonicity, decisiveness, and Pareto viability) operationalize this form of moral competence so it can be assessed from behaviour alone, without reference to a moral standard or expert baseline, forming a structural floor for alignment rather than a normative target. Testing nine frontier models on three simulated agent deployments with moral dilemmas, under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, no model expresses a coherent policy across the three deployments: surface-form paraphrase alone shifts verdict rates by up to 99 percentage points at a single escalation level, and competence on one scenario does not predict competence on another. The authors conclude that LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.
TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors
Most LLM safety benchmarks reduce responses to a binary refuse-or-comply score and ignore how implicit the threat in a prompt is. TIER spans four risk domains and four threat levels from explicit harmful requests to sophisticated jailbreaks, grading responses on a six-label behavior scale with two independent LLM judges. On six open-weight models, safety behavior shifts gradually across threat levels rather than flipping from refusal to compliance; contextual prompts produce the most varied behaviors, jailbreaks expose the largest robustness gaps, and models with similar attack success rates can have markedly different response distributions.
Uncensored Open-weight Models: Redistribution as the Persistence Layer
A growing set of actors strips safety guardrails from open-weight models and redistributes the results, and this study maps who produces them, who repackages them, and what gets built on top. Between January 2024 and March 2026 the authors identified 3,471 original uncensored models on HuggingFace, each repackaged an average of 2.4 times, with three actors responsible for 52% of the 8,164 compressed redistributions. Once quantized and mirrored across accounts, formats, and registries such as Ollama, the models persist regardless of upstream takedowns and become easier to deploy; of 1,643 GitHub applications integrating uncensored large language models (ULLMs), 25% were classified as explicitly malicious.
Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
Large language model (LLM) decision components inside agent workflows often return a judgement together with the factors that supposedly drove it, and operators may use those factors to monitor, diagnose, or escalate outputs, which only works if the explanations agree with the model's observable behaviour. The authors test two readings of a cited factor, necessity (changing it changes the output) and sufficiency (keeping it while removing other changeable information preserves the output), using controlled black-box interventions on two synthetic tasks: recommending advisors to clients and judging prompts for harmfulness or risk. Across eight models from the Claude, GPT, and Gemini families, mean Spearman correlations between the cited ranking and the measured necessity and sufficiency scores are 0.349 and 0.354 for advisor recommendation and 0.431 and 0.580 for prompt monitoring. An uncited factor outscores the lowest-scoring cited factor in about 58% of advisor responses, so the cited top three carry useful information but do not reliably identify the factors with the strongest measured influence.
5 more specialized papers
- Auditing Bias and Safety in Voice AI Customer Care Vignesh Ethiraj, Ashwath David
- Client-Side Probing of Deleted Ridge Statistics in Federated Unlearning Yijun Quan, Giovanni Montana
- Leveraging Imperfect Restoration for Data Availability Attack Yi Huang, Jeremy Styborski, Mingzhi Lyu et al.
- How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions Fernanda Mansilla, Aloysius Tok, Bahia Guella\"i et al.
- Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny? Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert
Theory 19
Towards a universal language of concepts: A survey
Humans learn and generalize new concepts from sparse data because they represent knowledge in rich structured formats, and the authors argue that programs are a strong candidate for a universal representation of concepts. The survey reviews computational models of concept learning that use programs as their concept representation and evaluates how far each contributes toward a universal representational language.
Learning-Augmented Algorithms: Guarantees, Construction Mechanisms, and System-Level Implications
Learning-augmented algorithms exploit machine-learned predictions that may be wrong while keeping formal worst-case guarantees. This survey synthesizes prediction interfaces, error measures, consistency-robustness trade-offs, and five representative construction mechanisms across online optimization, caching, learned data structures, graph problems, and mechanism design, and adds a theorem-level axis that separates achieved upper bounds from matched asymptotic dependence. It keeps formal guarantees distinct from empirical systems evidence, treats prediction cost, feedback, and composition explicitly, and lays out open problems in cost-aware prediction, endogenous error, semantic predictors, and benchmarking.
Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens
Supervised fine-tuning (SFT), few-shot in-context learning (ICL), KL-regularized reinforcement learning from human feedback and from verifiable rewards (RLHF/RLVR), on-policy distillation, and test-time search are usually treated as distinct paradigms, which makes results like the mixed effect of few-shot prompting on RL-tuned reasoning models look puzzling. The note frames them all with a two-step template: build a generalized Bayes or Gibbs posterior over outputs from a reference model and a utility signal, then approximate it by a forward-KL projection onto a parametric family, either in-weights or in-context. KL-regularized RLHF/RLVR, reward-weighted SFT, reward-weighted ICL, and advantage-weighted SFT all emerge as forward-KL projections onto reward- or advantage-induced posteriors, with the equivalences holding for objectives and first-order updates but breaking on the source and granularity of the learning signal. The framing also explains why supervised warm-up is practically unavoidable for importance-weighted projections and casts DeepSeek-R1 and o1-style models as test-time Bayesian search combined with training-time KL amortization.
Optimal Rates for Agentic Networked Information Aggregation
In a networked learning model introduced by Kearns, Roth, and Ryu, agents arranged in a directed acyclic graph each see only a subset of the raw features plus their parents' predictions, fit a linear predictor, and pass only their own prediction forward, mirroring how agentic AI systems propagate conclusions rather than data. Prior work showed the last agent's excess mean squared error on an M-covered path of depth D is O(M/√D), with a lower bound of Ω(M/D) for D < M², leaving a gap. The authors close this gap: the excess error is constant up to depth M² and Θ(M²/D) beyond it, via a sharper analysis of the cyclic instance and a new construction achieving Ω(M²/D) at every depth D ≥ M². They also show that for any fixed distribution the excess error contracts geometrically along the path, and prove the same optimal rate for logistic classification in the logit-passing model with binary cross-entropy (BCE) loss.
15 more specialized papers
- Data-Driven Learning of Unknown Nonlinear Differential Equations Using Functional Analysis Seyyed Shaho Alaviani, Yongzhi Qu, Gregory W. Vogl
- On the Abundance of Critical Points of the t-SNE Energy Nakul Haridas, Ryan Murray
- Nested Inductive Bias Framework for SPD Manifold Learning Tushar Das
- Centered Permutation Prefixes for SGD with Random Reshuffling: Sharp Rates, H\"older Geometry, and Composite Proximal Extensions Jiaxiang Li
- Representation Redundancy and Structural Complexity in Finite-Field Inversion Zheng Zhang, Na Zhang
- Interpretability for Turing Machines Billy Snikkers, Rumi Salazar, Daniel Murfet et al.
- Simulation-free Unbalanced Dynamic Optimal Transport with General Growth Penalty Junda Ying, Yuxuan Wang, Bowen Yang et al.
- CPR-IE:A Compression-Prediction-Resource Intelligence Efficiency Metric Xiantao Jiang
- Minimax Lower Bound for Estimating Diffusion-based Local Intrinsic Dimension Jaehee Seo, Wontae Jeong, Jisu Kim
- Solving Hard XAI Queries Based on a Compiled Dual-Rail Encoding Arthur Ledaguenel, Florent Capelli, Jean-Marie Lagniez
- An Analysis of Self-supervised Pre-training with Dependent Samples Maximilian Fleissner, Debarghya Ghoshdastidar, Samory Kpotufe
- Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy Margherita Mele, Andrea Castagna, Roberto Menichetti et al.
- PAC-Bayesian Reconstruction Guarantees for Time Series Variational Autoencoders Chlo\'e Hashimoto-Cullen, Ghislain Agoua, Benjamin Guedj et al.
- Dimension-Adaptive Batched Lipschitz Narrowing Without Knowing the Zooming Dimension Yasong Feng
- Shallow neural network approximation in mixed Sobolev spaces Yuwen Li, Guozhi Zhang
Multimodal 17
GEPARD - Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue
Real-time spoken dialogue needs text-to-speech that streams audio as text arrives, ideally served by an unmodified standard LLM engine. GEPARD is a decoder-only transformer trained jointly on text and audio embeddings and decoded to waveform through an FSQ-based neural codec, with auxiliary mechanisms such as zero-shot voice cloning, text augmentation, and classifier-free guidance moved out of the autoregressive decode loop into prefill or distilled into the weights so the backbone runs on vLLM without custom kernels. A single stream reaches a real-time factor of about 0.067, and 256 concurrent streams reach an aggregate speedup of about 204x on one server-class GPU. The report also diagnoses a short-register failure mode where autoregressive speech decoders break on one- or two-word inputs and distills two-pass classifier-free guidance into single-pass weights via Direct Preference Optimization (DPO).
MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering
Medical visual question answering (Med-VQA) is commonly assumed to need medical fine-tuning, large models, or multi-agent pipelines. MedProb is a lightweight probing framework that predicts multiple-choice Med-VQA answers directly from frozen vision-language model (VLM) representations without free-text generation. On PATH-VQA, SLAKE, and VQA-RAD it recovers substantially more answer-relevant signal than prompting and outperforms medical VLMs and agentic systems, and probing narrows the apparent gap between small and large models, suggesting smaller VLMs hold more recoverable signal than generation-based evaluation reveals. Across 14 matched general-purpose and medical VLM pairs, medical adaptation does not consistently improve linear decodability, free-text generation shows an answer-position bias of up to 10 percentage points, and the probe extends to open-ended generation through a rejection-sampling scoring procedure.
Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs
Growing chest radiograph volumes create a triage bottleneck, and most existing AI tools are unimodal binary classifiers with no notion of severity. The cross-modal triage network (CMTN) fuses a Swin Transformer V2 image encoder with a PubMedBERT text encoder through gated cross-attention, trained on 34,639 image-text pairs from MIMIC-CXR-JPG with an ordinal focal loss for four-tier severity and binary cross-entropy for 14 pathologies. Against reference labels it reaches a quadratic weighted kappa of 0.934 and macro-AUROC of 0.997 at 34 ms latency, well above the BioViL baseline, but a blinded audit against a radiologist's severity judgments yields a kappa of only 0.14, and only 54.3% of attention heatmaps localize acceptably. The gap shows that benchmark performance against NLP-derived labels does not establish clinical readiness without radiologist-labeled ground truth.
Knowing When Not to Answer: Pseudo-Ensembles for Abstention in Music Audio-Language Models
Multiple-choice accuracy cannot separate a lucky guess from genuine musical understanding, and full ensembles are too costly for music audio-language models, leaving single-pass entropy as the only confidence signal. The approach builds pseudo-ensembles from one pretrained model by applying answer-preserving perturbations, chiefly shuffling the order of candidate options but also corrupting the audio or swapping option labels, then averaging the predictive distributions so that ensemble uncertainty measures such as mutual information become available. On MuChoMusic with TinyMU, averaging over four option orderings raises accuracy from 55.7% to 59.2% and lowers the area under the error retention curve from 0.293 to 0.261 relative to single-pass entropy. The cost is a few extra forward passes with no retraining.
Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents
Embodied agents operating over hours or days need a memory in which the movements of dynamic objects can be queried in natural language, which video-language embeddings, geometric SLAM, and task-focused working memories do not provide. Linguistic Trajectory Encoding (LTE) compresses each object's motion history into a hybrid of natural-language descriptions, sparse spatial anchors, and visual anchors, adapting the compression to motion complexity and anchoring unobserved periods to the last seen location. On the new Spatial Memory Benchmark built from EgoLife multi-day recordings, it reaches 45.3% success on semantic trajectory retrieval and 48.7% on long-horizon object retrieval versus 31.9% and 34.4% for the best prior method, with 8.7x to 26.1x compression and sub-second queries over 24 hours of video, and it also outperforms EgoVLPv2 on Ego4D natural-language queries.
PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation
Evaluation of text-to-audio-video (T2AV) generation tends to fold audio into overall video quality or score it apart from any audiovisual grounding, which hides where systems actually fail. PRISM-Bench draws on 900 human-verified samples and factors audio evaluation along two axes, audio type (speech, music, sound) and whether the sound source is visible on screen, scoring four perceptual dimensions through 35 fine-grained criteria with a multimodal-LLM judge doing blind side-by-side comparison against ground-truth references. That protocol agrees with human raters over 70% of the time on average, and applying it to recent systems exposes a large gap between frontier and open-source models plus a shared tendency to optimize perceptual fidelity while failing at grounding and control, worst for music and synchronized on-screen audio.
MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression
Multimodal reasoning models produce long chains of thought (M-CoT) that inflate compute and key-value cache pressure, and existing compression methods lacking cross-modal constraints tend to induce visual laziness and hallucinated reasoning. Modality-Contrastive Preference Optimization (MCPO) is a two-stage method needing fewer than 900 training samples: a step-level Normalized Cross-Modal Mutual Information (NCMI) pruning algorithm compares reasoning with and without the image to remove steps that do not depend on the visual input, then supervised fine-tuning followed by an asymmetric length-controlled preference loss enforces shorter trajectories in the with-image context while preserving modality consistency in the no-image context. On base models such as Qwen3-VL-Thinking it shortens chain-of-thought length by up to 69.5% and speeds end-to-end inference up to 3.34x while preserving original accuracy.
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Scientific papers demand reasoning across text, equations, figures, tables, code, and datasets while keeping track of where supporting evidence came from, yet existing benchmarks test these skills in isolation. SciDocBench contains 124 expert-authored, difficulty-screened questions in seven research-assistant capability groups and 19 subtasks across five scientific domains, each instantiated in English and Chinese with images-first or interleaved document layouts for 496 controlled evaluation instances. The strongest evaluated system scores only 62.6/100, with pronounced weaknesses in document perception, evidence grounding, verification, and cross-document reasoning. To turn these diagnostics into training data, the authors add SciDocIR, a typed evidence-graph representation preserving document objects, layout, cross-references, and provenance, and use it to build SciDocDataset with roughly 15K supervised fine-tuning and 8K reinforcement-learning samples over 14 verifiable subtasks.
From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making
Vision-Language Models (VLMs) are usually judged by their final answers, which leaves open whether those answers actually rest on visual evidence. The authors apply layer-wise causal interventions on video-text attention pathways in a video-based generative multiple-choice setting, targeting spatial, causal, and temporal reasoning. Visual information is integrated mainly while the model processes the candidate answer options, which act as the primary textual grounding sites for the decision; nouns serve as semantic anchors during multimodal enrichment, while verbs matter more when temporal relations are processed. A distinct pattern in temporal questions suggests VLMs struggle to reconstruct event order across frames, though the authors note this fragility may partly reflect linguistic biases in the temporal expressions used to define relations between events.
8 more specialized papers
- TRILOGUE: A Trilingual Spoken Dialogue Fact-Checking Benchmark with Evidence and Paired Audio Chaewan Chun, Meruyert Aristombayeva, Jiyoung Choi et al.
- Tracing Audio Grounding and Answer Selection in Audio LLMs Hyebin Cho, Suho Yoo, Jihoo Jung et al.
- Latent-Aligned Reasoning for Multimodal Recommendation Jiarui Jin, Anyang Ji
- Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion Xu Lin, Ke Wang, Hui Kang et al.
- Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models Minji Kim, Jihyoung Jang, Hyounghun Kim
- Whose record is this? Diagnosing and authorizing record use in personalized multimodal models Xinyu Mao, Junsi Li, Chenyang Liu et al.
- MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain Sourav Malakar, Harshit Nigam, Akash Ghosh et al.
- MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models Changming Xiao, Zhenliang Ni, Jinhui He et al.
Vision 15
Where Appearance Fails, Geometry Recognizes: A CAD-Free 3D Shape Prior That Complements Vision Foundation Models
Recognizing specific objects onboarded without labeled training data is common in manufacturing and service robotics, but the usual renderable prior, a computer-aided-design (CAD) model, is often unavailable, and frozen foundation features struggle on low-texture, geometrically similar parts. Each object is instead reconstructed from a short scan with 3D Gaussian Splatting (3DGS), summarized into a per-class shape prototype, and fused with frozen DINOv2 image features. Geometry from RGB-D depth, 3DGS, and CAD gives comparable recognition, and on shape-distinctive household objects in HOPE geometry alone reaches 0.920 versus 0.832 for images alone, while on textureless industrial parts in T-LESS fusion lifts accuracy from 0.560 to 0.591. The prior rescues more image failures than it breaks, helps most under partial occlusion, and its value lies in geometry rather than pixels, since 3DGS renderings do not improve the image side.
What Moves? Localized Motion Representations for Compositional Scene Control
Most video representations encode motion globally, and localized embeddings computed from crops or post-hoc feature masking discard the camera motion and scene layout needed to interpret an entity's movement. The proposed promptable localized motion representation processes the full video while conditioning the motion encoding on a user-specified spatial mask, producing temporally consistent, region-addressable embeddings that isolate local dynamics while keeping global context for disambiguation. The embeddings enable object-level motion transfer for controlled composition of dynamic scenes and support localized action classification in multi-actor videos, outperforming global representations localized by cropping or masking on both tasks.
LookThere! Sparse Vision by Reinforced Selection
Vision transformers process every image token even though most tasks need only a small fraction, and existing token-selection methods break down at extreme sparsity or rely on heuristics such as token diversity and attention scores. LookThere jointly trains a shallow input selector and a deep representation extractor end to end with reinforcement learning, so the selector learns where to look and the extractor learns what to see without auxiliary signals. It maintains accuracy on high-resolution sparse recognition tasks such as traffic signs and billiards using as little as 0.2% of the input, and surpasses prior selection methods across ImageNet classification, ADE20K segmentation, zero-shot classification by distillation, and counting regression.
Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions
Conventional convolutional vision models identify objects in isolation, whereas commonsense knowledge lets a system interpret relationships among objects, actions, and context in everyday scenes. The survey systematically reviews how commonsense is injected into computer vision through knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers. It catalogs open problems around dataset bias, incomplete knowledge sources, and integration difficulty, and points to cross-modal reasoning, scalable knowledge injection, and hybrid neuro-symbolic architectures as directions forward.
Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions
Chilli is an economically important crop in India whose diseases are hard to identify without experts, and while Vision Transformers (ViTs) classify them accurately, their footprint hinders on-device deployment. The authors propose a unified compression pipeline combining Hessian-Balanced Adaptive Block Pruning (H-BAC), guided by second-order sensitivity, with quantization and attention-based knowledge distillation, first ablating each technique independently and then integrating the best components. On a three-class chilli dataset with a genuine cross-village, cross-device out-of-distribution test split, compressed models match or exceed the 95.13% FP32 baseline with 74 to 98% size reduction, and the fully integrated pipeline shrinks the model 54.5x, from 327.42 MB to 6.01 MB, at 95.13% accuracy. A directly trained student of the same 6.01 MB INT8 size reaches a comparable 94.87% without pruning or distillation, marking where the added machinery is and is not yet shown to be worth its cost.
10 more specialized papers
- The microscope is the mask: privileged views and labels from a cryo-ET forward model Bogdan Toader, Kiarash Jamali, Tanmay A. M. Bharat et al.
- Dual-Part Multi-Lateral Branched Network for Multi-Class Segmentation in Cardiovascular Catheterization Angiograms Olatunji Omisore, Ahmed Elazab, Ali Shahidinejad et al.
- Mitigating Performance Discrepancy in Cross-Domain 3D Class-Incremental Learning Jinge Ma, Gautham Vinod, Bruce Coburn et al.
- SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection Yongchun Lin, Xinliang Zhang, Yun Zou et al.
- Sound-based Multi-Person 3D Pose Estimation Yusuke Oumi, Yuto Shibata, Go Irie et al.
- VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition Jiangang Zhu, Zheng Wang, Bin Zhu et al.
- Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection Changyi Li, Yu Xiao
- Adaptive Gated Deepfake Detection for Low-Resolution and Resource-Constrained Environments Vaishnavi Sen, Cody Laurie, Rashida Hasan
- Reflection-aware Generative Novel View Synthesis GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh
- UniMate: One Unified Model to Animate Diverse Skeletons Linzhan Mou, Jiahui Lei, Zhiyang Dou et al.
Reasoning 11
Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning
Humans who can solve 2+5 can also solve 'two plus five', but LLMs are far less accurate on verbal renditions of arithmetic problems they solve almost perfectly in numeric form. Using attribution patching, the authors localize the circuit each model uses for numeric arithmetic and for verbal problems in English, Spanish, and Italian, then test whether overlap with the model's own numeric circuit predicts how well it generalizes to each verbal format. Circuit overlap accounts for the relative difficulty of the three verbal formats, which models generalize best, and which individual items are solved correctly, rivaling supervised probes while requiring no labeled data.
Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective
Pause tokens inserted into sequences improve LLM reasoning, and prior explanations focus on added computational expressivity rather than on how they change fine-tuning dynamics. Two controlled pilots reveal asymmetries: on a synthetic continual-learning task, masked pauses overwrite a previously learned distribution roughly 4x less at matched adaptation, which the authors call mode retention, and on a synthetic math probe the token adjacent to a step boundary comes to encode substantially more downstream-step information, called non-myopic compression. These motivate Masked Boundary Pause (MBP), which places pause tokens at reasoning-step boundaries with their loss masked. Across 1B-8B Qwen and Llama models, MBP improves reasoning by up to 6 points on math and 2.5 points on code while preserving general language understanding, and the gains carry over to GRPO training.
Extremely Sparse Supervision Incentivizes Reasoning Ability
Prevailing post-training methods for reasoning optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. Working in the on-policy distillation setting with the Qwen3 family, the authors find that supervising as few as one or two tokens per reasoning trajectory, about 0.05% of all tokens, matches or surpasses full-token training on mathematical reasoning in most cases. The effect holds across nine teacher-student configurations of varying scale and is further validated on coding reasoning, Llama models, and Proximal Policy Optimization (PPO)-based reinforcement learning with verifiable reward (RLVR). The authors suggest this mirrors natural learning, where one reflects on a few critical steps rather than correcting every word, and argue it points toward more efficient post-training algorithms.
ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying
Group Relative Policy Optimization (GRPO) and related reinforcement learning (RL) methods train large language models (LLMs) to reason using only final-answer rewards, which say nothing about which intermediate steps helped or hurt and grow increasingly sparse as reasoning chains lengthen. ConsensusBench supplies rule-based process-level signals by sampling N rollouts, keeping the correct ones, and clustering semantically equivalent intermediate statements into Consensus Nodes, verifiable sub-outcomes that correct answers tend to pass through; a process reward derived from these nodes, ConsensusPR, is added to GRPO-style training, and three metrics (Final Answer Accuracy, Node Coverage Rate, and Tokens per Node) support process-level evaluation. Across AIME 2024, AIME 2025, GSM8K, MATH-500, and ConsensusBench, the method consistently outperforms GRPO-style baselines.
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
Chains of thought contain distinct reasoning operations such as problem formulation, goal decomposition, and deduction, but it has been unclear whether these operations correspond to any structure in a model's hidden representations. The authors probe held-out activations and find that reasoning operations are separable in representation space, with separability peaking in middle layers, and they rule out lexical and positional confounds as the explanation. Further analysis shows that identical surface tokens are represented differently depending on the operation of the surrounding chunk, and attention-masking interventions show that operation-aligned representations at the start of a chunk depend on the preceding reasoning context.
Fractal basins trap latent reasoning
Reasoning models take longer on harder problems, but the mechanism behind these slowdowns has been unclear. Treating reasoning traces as trajectories of a dynamical system, the authors find that leading reasoning models exhibit transient chaos and have fractal basins of attraction, with fractality increasing with task difficulty across Sudoku, maze solving, visual puzzles, and mathematical logic. The chaos arises because reasoning lingers near saddle points, which they show correspond to nearly-correct candidate solutions, implying that longer reasoning on hard problems is an inevitable consequence of problem hardness rather than a fixable inefficiency.
AxQM: A Textbook-Scale Benchmark for Formal Proof Synthesis in a Library of Finite-Dimensional Quantum Mechanics
Formalizing mathematics in a proof assistant sets a standard of rigor that physics rarely meets, since theoretical arguments carry unstated idealizations whose gaps can cascade through dependent results. AxQM provides 1,019 kernel-checkable proof-synthesis tasks over 479 items drawn from Nielsen and Chuang's Quantum Computation and Quantum Information, stated in a custom Lean library of finite-dimensional quantum mechanics. By task count it is the largest proof-synthesis benchmark in physics by a factor of four, and because it derives from a near-complete formalization of the textbook's formal portions, every task has a known solution that the authors keep private. Grading is deterministic: the Lean kernel checks that a proof compiles, contains no sorry in it or its dependencies, and introduces no new axioms.
Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory
Knowledge Space Theory (KST) formalizes the idea that mastering a concept requires first mastering its prerequisites, and the question is whether LLMs' mathematical competence is organized the same way. The authors build a KST-grounded evaluation that treats principled knowledge dependencies as a norm and compare eight open- and closed-source models against real human learners. LLMs frequently violate prerequisite dependencies and fail to use related knowledge supplied in context to solve dependent questions, and different models show low overlap in their knowledge distributions, so they do not even share a structure among themselves. These deficits stay largely invisible to accuracy-based and LLM-as-judge evaluations.
3 more specialized papers
- DODR: Deterministic Operator-Driven Reasoning in Latent Space Weicai Huang (Beijing MQPat Technologies, Co., Ltd.)
- A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR Thi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo et al.
- GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity Shuang Liang, Xin-Yu Hu, Xiang-Jun Ou et al.
Robotics 9
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
Pretrained vision-language-action (VLA) models handle broad manipulation but remain unreliable on tasks demanding precision and repeatability, and applying real-world online reinforcement learning (RL) to improve them runs into unreliable value signals that cause policy drift and large-model overhead that limits throughput. VLA-Precision addresses both with the Asymmetric Co-Bootstrapping (ACoB) algorithm, in which early intervention-guided behavioral learning rapidly lifts performance and experience quality while global return propagation and local preference ranking progressively calibrate value estimates for reference-regularized policy improvement, and with ACoB-Stream, a closed-loop experience-policy architecture built on invariant-state decoupling and on-demand streaming that delivers up to 10.9 times higher throughput and computational efficiency. Across nine high-precision chemistry tasks spanning four categories and four robot embodiments, it reaches a 98.3% mean success rate in 45.8 minutes of training per task, with 27.6-second episodes running at 1.2 and 1.8 times the speeds of VLA and RL baselines.
Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI
Machines sent into dangerous, mission-critical settings must learn new skills after deployment from scarce data on onboard compute without forgetting prior competence. CFAM pairs a frozen slow-learning component made of Sensor, Reasoning, and Action cortices with a fast-learning Capsule Field that stores field experience one-shot and gradient-free as Competence Capsules, so skills are installed few-shot in the lab and extended continually in the field. Evaluated on manipulator, quadruped, humanoid, quadrotor, and off-road vehicle embodiments against pi0, CogACT, and SpatialVLA, it matches a standard policy trained on the full dataset using 40% of the data, or 2.5x fewer trajectories. Autonomous capture of verified near-out-of-distribution cases raises action success by 13.9 percentage points, and backward transfer in sequential simulation is -0.5 points versus -11.4 for LoRA.
One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation
A single pretrained diffusion model of traffic can act both as an ego motion planner and as a controllable generator of safety-critical scenarios for stress-testing planners. On the planning side the authors introduce a Single-Stream Dual-Stream (SSDS) diffusion-transformer decoder that fuses scene context through joint attention instead of late cross-attention, plus Decoupled Annealing Posterior Sampling with Energy (DAPSE), a training-free guidance scheme that applies arbitrary energy functions at the clean-sample level without auxiliary networks. Inference-time guidance then steers selected agents toward aggressive cut-ins, lead-vehicle braking, and combined longitudinal-lateral maneuvers while keeping surrounding traffic realistic. In closed-loop nuPlan simulations against independent black-box planners, the generated scenarios expose failures hidden by standard benchmarks, and the SSDS planner, despite stronger nominal scores, degrades more under these scenarios, showing that benchmark superiority does not guarantee robustness.
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Existing Vision-Language-Action (VLA) benchmarks mostly measure task completion in predefined settings and reveal little about how models reason as spatial and procedural difficulty increases. RoboSPA (Robot Spatial-Procedural Assessment) is a large-scale manipulation dataset and benchmark targeting fine-grained spatial reasoning and long-horizon procedural planning, with 10 task categories and 56 base tasks each instantiated at five difficulty levels for 280 variants, plus 527K trajectories collected across multiple embodiments and diverse scenes. Beyond binary success rate, it adds diagnostic metrics for finer-grained evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning.
Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models
Vision-language models (VLMs) are increasingly used as reward functions for robot learning, which requires paraphrase invariance: the same trajectory should receive the same reward under semantically equivalent goal descriptions. ROBORMBENCH measures this with 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases spanning lexical, syntactic, and action-goal rewrites. Across proprietary and open-source VLMs, paraphrasing the instruction alone substantially shifts predicted progress scores and can flip identical robot behaviour between failure and success, with instability growing under more divergent rewrites and not reliably reduced by scale or explicit reasoning. Dedicated reward models trained with trajectory-grounded supervision are substantially more stable, which the authors present as evidence that paraphrase robustness is a core requirement for VLM-based reward modeling in robotics.
4 more specialized papers
- Modular Deep Recurrent Neural Network: Application to Quadrotors Nima Mohajerin, Steven L. Waslander
- Coupled Control and Wireless World Models for Resilient Remote Robotic Control H. P. Madushanka, Sumudu Samarakoon, Mehdi Bennis
- A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning Chongwen Dong, Mithun Paul Saint-Germain, Pinjari Asif et al.
- What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies Vivek Chavan, Pengtao Xie, Yahuan Shi et al.
Reinforcement Learning 4
Spectral-Target Physical Latent Structuring for JEPA-Style World Models
Latent world models such as LeWorldModel (LeWM) jointly train encoder and predictor with regularizers like SIGReg to prevent collapse, yet in highly dynamic environments their latents can suffer a distinct failure the authors call physical representation laziness: the states do not collapse but fail to encode key physical properties, causing widespread planning failure. The fix is a lightweight Fourier auxiliary head that supervises the latent space toward physically informed spectral targets during training, adding no inference-time cost and applying to any environment. Planning success rates improve substantially in dynamic environments where the baseline exhibits laziness, with modest gains elsewhere and particularly large benefits in low-data regimes. Higher latent correlations with physical properties accompany the better planning, supporting the link between physically structured latents and downstream performance.
SQL-Zero: Self-Evolving Text-to-SQL
Training Text-to-SQL models normally depends on expensive, domain-specific human-annotated question and SQL pairs. SQL-Zero instead runs proposer-solver self-play where a challenger and a solver start from the same base model, the challenger generates SQL pairs calibrated to be hard but solvable at the solver's current level, both roles are updated with Group Relative Policy Optimization (GRPO) in alternating turns, and a template-level repetition penalty keeps the challenger from collapsing in diversity, with database execution as the only ground truth. Training label-free on BIRD databases improves over the zero-shot base on BIRD dev by 6.6 points at 3B and 7.3 points at 7B, and scores above a matched control trained on human gold labels, though transfer to unseen Spider databases and lexically perturbed Spider-Syn holds across every iteration only at 3B, while at 7B only the first iteration preserves it.
Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning
Cooperative multi-agent reinforcement learning (MARL) agents learn from past experience that can become unreliable if the environment or task objective shifts mid-training, so agents first need to detect that a change has occurred. Patterns of Past Rewards (PPR) is a lightweight, algorithm-agnostic detector that smooths agents' return streams, emphasizes recent changes, and applies a statistical drift detector to flag significant shifts. Evaluated in a custom Speaker-Listener environment built on the Multi-Agent Particle Environment under two controlled non-stationarity scenarios, the method exposes a trade-off between detection speed and alarm stability. A smoothed-return baseline detects shifts earlier but fires many repeated alarms, while running the detector on raw returns often misses the shift entirely; PPR limits redundant detections while still catching the controlled changes.
1 more specialized paper
- Compact Bellman-Grounded Cognitive Maps for Cost-Aware Navigation Yuzhe Han, Mingkun Xu, Yujie Wu