Monday, August 17, 2026
Highlights
MobileMem: Learning from a Year of Mobile Experiences
Personal assistants that accumulate knowledge about a user over months need benchmarks built from heterogeneous, multimodal, evolving personal data, which existing long-term memory evaluations do not provide. MobileMem supplies both a benchmark and framework for on-device memory, using a knowledge-grounded synthesis pipeline to turn a year-scale collection of real user-app sessions into coherent, temporally consistent long-horizon trajectories. It offers matched text-only and multimodal settings that probe multi-hop and temporal reasoning, updating superseded knowledge, and inferring preferences the user never stated, framing memory as accumulated experience rather than a retrieval index over isolated facts.
Persistent mobile assistants need memory spanning a year of heterogeneous, cross-app, multimodal activity, but existing long-term-memory benchmarks assume cloud storage and clean dialogue histories. MobileMem synthesizes year-scale interaction trajectories from real smartphone usage metadata and interviewed personas, then evaluates nine memory systems end-to-end on multi-hop, temporal, preference, and unanswerable questions.
- The
KEMEsynthesis engine alternates top-down temporal planning — recursively expanding a persona and time horizon into an event graph, then into sessions — with bottom-up experience evolution that revises unexpanded future events and updates persona attributes with message-level evidence links, using realOPPOapp-usage metadata as immutable "knowledge anchors"; the text track covers seven app templates whileMobileMem-Omniadds screenshots, synthetic portraits, and bilingual dialogue in trajectories exceeding 2M tokens. - Systems that preserve raw conversational detail win:
A-MEMscores 79.68 overall andHippoRAG278.85 with aGPT-4.1-minibackbone (HippoRAG2reaching 80.06 onGPT-5.4-mini), while a 128k-token long-context baseline reaches only 54.51 and compress-and-overwrite designs likeMem0(35.63) andLangMem(24.79) collapse. - Memory construction cost varies by roughly 4x for similar accuracy —
A-MEMspends 5.46M tokens per trajectory onGPT-4.1-miniand 11.17M onGPT-5.4-mini(which extracts far more keywords per memory unit), whereasHippoRAG2matches or beats it at about 2.8M. - Temporal reasoning is the consistent weak spot, topping out at 72.04 against 86.30 on single-hop, and adversarial unanswerable questions invert the whole ranking:
LangMem, last overall, leads at 76.42 because stronger retrievers surface weakly relevant distractors that convince the model an answer exists. - The main caveats are that trajectories are LLM-synthesized by
GPT-5.2from only two volunteers in the text track and eight real plus eight virtual personas inOmni, global consistency is checked by manual sampling rather than any automatic metric, only about 44% of fine-grained profile fields ever surface in the generated data, and — despite the on-device framing — every evaluated system runs cloud-hosted backbones with cost reported purely in tokens, never in storage, latency, or power.
Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning
A hidden Markov model (HMM) does three things — infer a belief over hidden state, propagate it through a transition, and emit back into observation space — and the argument here is that time-indexed Predictive Information Bottleneck VJEPA (PIB-VJEPA) has exactly that structure, with the stochastic context encoder acting as an amortized filter, the probabilistic predictor as latent dynamics, and the decoder or inverse target encoder as the emission. Four progressively stronger levels of correspondence are defined along with sufficient conditions for exact sequence-level HMM equivalence, and MCJEPA makes the link concrete by replacing the latent predictor with a learned transition matrix whose powers guarantee exact multi-horizon Chapman-Kolmogorov consistency in the finite time-homogeneous case. Deterministic temporal JEPA falls out as a degenerate Dirac-kernel case, and controlled experiments confirm transition composition, the filtering interpretation, and predictive Markovization on a synthetic process with known structure.
Predictive-coding architectures like JEPA are usually described in terms of representation matching rather than probabilistic inference, leaving unclear what generative model, if any, they correspond to. The claim here is that a fully stochastic temporal JEPA — specifically PIB-VJEPA, where the current representation, the future target, and the transition are all distributions — instantiates the encode–transition–emit pipeline of a hidden Markov model, with the context encoder acting as an amortized filtering distribution, the predictor as the latent transition, and a decoder or inverse target encoder as the emission.
- The correspondence is graded rather than binary, split into four levels — shared computational roles, an emission-complete latent state, sequence-level HMM equivalence, and model-and-objective equivalence — with a theorem giving four sufficient conditions for the third level: first-order Markov latent dynamics, a valid state-to-observation conditional, transition-consistent latent marginals, and a history encoder that coincides with the induced filtering posterior.
MCJEPA(Markov-Chain JEPA) makes the idea concrete by replacing the neural predictor with a learned row-stochastic transition matrixAoverKcategorical states, so an h-step forecast is justq_t A^hand matrix powers give exact Chapman–Kolmogorov path consistency — every decomposition of the same horizon yields an identical predictive distribution, which is what makes hierarchical planning over composed short- and long-horizon steps well-defined.- Because the encoder, target encoder, and transition are trained jointly against a KL latent-matching loss, they can agree through degenerate solutions, so training adds an occupancy term
KL(q̄_B || Unif(K))against single-state collapse plus a per-sample entropy penalty against uniform-assignment collapse — two regularizers whose weights must be balanced, since too much occupancy pressure forces artificial state usage and too much entropy pressure causes premature hard assignments. - The construction generalizes cleanly along both axes of latent Markov dynamics: side-information-conditioned matrices
A_φ(ξ_t)for action- or time-dependent transitions, continuous-state Gaussian or flow-based kernels, continuous-time generators viaexp(Δt Q_φ)for irregular sampling, and deterministic temporal JEPA recovered as the degenerate Dirac-kernel limit as the transition covariance goes to zero. - The honest limits are substantial: standard JEPA training optimizes latent target matching plus information-bottleneck regularization, not observation-sequence likelihood, so the stated conditions are sufficient but not necessary and are not guaranteed by ordinary JEPA training; compressive target encoders are generally not invertible, the implicit Bayes-rule emission
p(x|z) ∝ p_data(x) q_θ(z|x)is a static correspondence that need not be tractable for generation; and the empirical support comes from 4 controlled experiments on synthetic processes rather than any real video or large-scale benchmark.
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
Agents in the ReAct paradigm think only during the Thought phase and sit idle while an action is serialized and the environment responds, leaving a recurring reasoning idle window unused. Second Thought is a training-free inference framework that forks four auxiliary reasoning branches the moment a Thought phase ends, decodes them concurrently with the main loop, and merges their output back when the observation arrives, moving the extra reasoning off the sequential critical path. Across three agentic benchmarks and three reasoning models it lowered average turn count in all nine model-benchmark pairs and cut main-thread decoding by up to 43% with Pass@1 unchanged in seven pairs; against a compute-matched control that spends the same budget on the main thread, it achieved strictly higher Pass@1 with 1.3 to 3.2 times less sequential decoding.
ReAct agents only reason during the Thought phase, leaving the Action–Observation interval — the wait for a tool to execute and return — as a recurring window in which no reasoning happens at all. Second Thought is a training-free inference framework that forks four auxiliary reasoning branches the moment a Thought concludes, decodes them concurrently with the main loop, and appends whatever they produced to the observation message so the next turn starts with the extra deliberation already in hand.
- The four branches each take a complementary angle on the same live trajectory —
Checkaudits the just-finalized plan for unverified assumptions,Recallresurfaces constraints fading from the context window,Rehearsepre-computes conditional next steps, andAlternativedrafts fallback strategies — so merging is plain concatenation rather than voting or ranking, and all four share the main thread's prompt prefix KV cache. - Because the window closes unpredictably when the observation arrives, every branch emits a stream of atomic thoughts: self-contained units of ≤25 words wrapped in
<thought>tags with no cross-references, so cancelling mid-generation discards at most the unit in flight, and each buffer is truncated at its last closed tag and capped at 5 thoughts per dimension. - Across
SWE-Bench Pro,Terminal-Bench 2.1, andτ³-benchwithDeepSeek-V4-Flash,Qwen3.6-Plus, andMiniMax-M3, turn count falls in all nine model–benchmark pairs and main-thread decoding in six by up to 43% (roughly 20% on average there), while Pass@1 is statistically unchanged in seven pairs and significantly higher in two (+12.4 and +10.2 points, both onTerminal-Bench 2.1); a paired wall-clock replay of 50SWE-Bench Proinstances turns this into 10.9% lower median per-task latency (256.9s to 229.0s), with branch contention costing only 2.8s against 27.9s saved. - Against
s1budget forcing — a compute-matched control that spends the same extra reasoning tokens on the main thread's own thought —Second Thoughtreaches strictly higher Pass@1 with 1.3× to 3.2× less sequential decoding in all four settings where the control is applicable, and a replay ablation shows harvested thoughts substitute for work the agent would otherwise do inline (next-turn reasoning grows from 196.2 to 316.5 tokens when they are removed). - The main costs are monetary and structural: four branches raise per-task API spend by 66.4% to 181.5% (almost entirely cached prefix reads, reducible to 16.3–35.5% by keeping only the
Alternativebranch), the default configuration is window-starved — letting branches run to completion on the critical path reaches 56.7% Pass@1 versus 52.0%, meaning the truncated version captures only 41% of the attainable gain — and onτ³-bench, where windows are short and failures stem from retrieval and policy adherence rather than planning, gains top out at +3.1 points.
From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models
A retrospective covering October 2018 to July 2026 tracks how language models went from BERT to agents that solve competition math and write software, reporting that the ability to resolve real coding issues improved nearly sixfold per year since late 2024 while cost per unit of capability collapsed, with budget-tier models matching flagship capability at one to six dollars per million tokens. It also documents fragmentation of the frontier into task-targeted models, with different systems leading on frontend coding, repository-level coding, and terminal tasks. A companion grade-school math experiment on Qwen2.5 shows basic decoding solving 58 of 100 problems against up to 79 with advanced sampling, and a confidence ranker placing 47 correct answers in its top 50, with all materials released publicly.
Eight years of language-model progress are measured against public benchmarks, API prices, and one reproducible small-model experiment, rather than against impressions. The picture that emerges is a capability curve still compounding fast, a cost curve collapsing at a similar rate, and a frontier that has fragmented into task-specific leaders — so the practical unit of deployment becomes a routing policy rather than a single model.
- A log-linear fit to
SWE-bench Verifiedrelease scores from October 2024 to July 2026 puts annual growth in the odds of resolving a real GitHub issue at ~5.8× (R²=0.78, n=14), carrying the independent frontier to 96–97% and implying the benchmark saturates within about a year, the same lifecycleGLUEandMMLUalready completed. - Input-token prices fell ~60× from
GPT-3's $60/M (2020) toGPT-5.6 Luna's $1/M (2026), and in OpenAI's own launch table Luna wins 7 of 10 agentic and professional-work benchmarks against the eleven-week-oldGPT-5.5flagship at one-fifth the price — while trailing sharply on long-context recall (41.3% vs 81.5% onMRCR) and hardest academic reasoning. - The mid tier is squeezed out: a Pareto analysis of fifteen
GPT-5.6effort settings finds everyTerraconfiguration dominated by a cheaper-or-smarterLunaorSolsetting, and effort itself has an interior optimum —Opus 5peaks onFrontier-Benchat xhigh (44.4%) and drops to 43.3% at max. - No model leads everywhere —
Opus 5tops frontend preference (1,712 Elo) andARC-AGI-3(30.2%, ~4× the prior 7.8% record),Fable 5leadsSWE-bench Proat 80.0% against Sol's 64.6%, andSolleads terminal work — yet a two-model router pairing Sol with Fable 5 recovers the entire +2.4-point gain of a six-model oracle on the fourteen-benchmark suite. - The inference-time experiment is honest about its own weakness: a frozen
Qwen2.5-1.5Bconfiguration on 100 lockedGSM8Kitems scores 58/100 greedy versus 62/100 for four-sample plurality (paired McNemar p=0.481, not significant) against a 79/100 any-sample oracle, so selection — not sampling — is the bottleneck; a post-hoc confidence model reaches 0.833 AUC and 47/50 correct in its top-ranked half, but it was designed after labels were visible and cross-validated on those same items. - Provenance is the standing caveat throughout: mid-2026 numbers mix vendor claims with independent harnesses (GPT-5.5 scores 85.1% vendor-reported versus 82.6% reconstructed), the oracle router cheats by knowing benchmark identity, and targeted gains may not transfer —
Opus 5's ARC-AGI-3 record collapses to a statistical tie on the held-outWitnesspuzzles.
Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation
The EU AI Act obliges providers of high-risk systems to document how their systems reach decisions, and circuit discovery in mechanistic interpretability is the natural source of such evidence — provided two competent analysts using the same tool with different defensible settings would file the same thing. A pre-registered grid crossed seven analytic axes, each level drawn from a published implementation, over GPT-2 small on the indirect object identification task, mapping every discovered circuit through a deterministic claim map into a structured Annex IV statement. Across 15,840 specifications, 7,561 of which produced a claim, the derived statement flips across 73.2% of specification pairs and the most common claim covers only 41.1% of the space; standardizing the evaluation metric leaves the flip rate at 59.4%, and dropping circuit size from the claim entirely still leaves 27.1%. The underlying circuits are near-disjoint (median pairwise Jaccard overlap 4%) and functionally uncorrelated (Cohen's kappa 0.015), so this is not one mechanism described in different words; the study covers one model and one task.
Circuit discovery is the most mature instrument mechanistic interpretability could offer to satisfy EU AI Act Annex IV documentation duties, but its output is only useful as evidence if two competent analysts using the same tool on the same system file compatible claims. A pre-registered multiverse over seven defensible analytic axes, with each discovered circuit mapped through deterministic code into a structured regulatory statement, finds that the filed claim flips across 73.2% of specification pairs.
- The design crosses seven axes — discovery objective, ablation operator, corruption distribution, evaluation metric, size threshold, prompt variant, seed and granularity — with every level taken from a published implementation, yielding 15,840 specifications on
GPT-2 smalland the indirect object identification task, of which 7,561 survived the metric-relative threshold rule and produced a claim. - The modal claim commands only 41.1% of the space, failing a proposed filability criterion (modal share ≥ 1 − α) at every tolerance a conformity body would plausibly accept, and the top two claims contradict rather than merely differ: 41.1% attribute the behaviour to early layers and 28.5% to late layers.
- No single control rescues the filing — standardising the evaluation metric, the most influential axis, leaves the flip rate at 59.4%, while removing circuit size from the claim entirely and holding size fixed still leaves 27.1% (95% CI 0.255 to 0.286).
- The instability is structural and functional, not verbal: median pairwise
Jaccardoverlap between circuits is 4% and per-example functional agreement is uncorrelated at Cohen's kappa 0.015, which directly rebuts the "phantom specialisation" objection that discovery merely samples behaviourally equivalent subgraphs. - Caveats are substantial and the authors foreground them — one model and one task (a
Pythia-160mreplication is pre-registered but unrun), a 52.3% discard rate skewed heavily by metric, one of the library's seven documented discovery objectives that crashes on its own canonical task, and a pooled-versus-within-size reversal showing discovered circuits are in fact more stable than a size-matched random null (0.2746 against 0.4230) once size is held constant.
Scaling Domain Data Repetition in LLM Pretraining
As models grow, token budgets grow with them to keep a sensible tokens-per-parameter ratio, but scarce high-quality domain data cannot scale the same way and gets diluted in the mixture; repeating it counteracts dilution at the risk of overfitting. Sweeping this trade-off under practical scaling, where token budget grows with model size, yields two findings: at fixed tokens-per-parameter the optimal repetition count mildly increases with model size, and across domains the optimal count correlates strongly and negatively with a domain's final validation loss, while the amount of unique domain data barely matters. The practical implication is that repetition counts tuned on small proxy models at the same tokens-per-parameter ratio carry over to larger runs.
Practical LLM scaling grows the training-token budget in step with model size, but high-quality domain data (code, math, curated text) cannot be scaled the same way, so its share of the mixture is progressively diluted. Repeating that scarce data counteracts the dilution at the cost of overfitting risk, and this work charts where the optimum sits as a function of model size, domain, and corpus size.
- The setup sweeps repetition counts for a fixed domain while holding the tokens-per-parameter ratio (
TPP) constant and scaling the token budget proportionally with parameters, using final per-domain validation loss as the criterion for the optimal number of passes. - Contrary to the usual intuition that larger models memorize faster and should therefore see repeated data less often, the optimal repetition count mildly increases with model size at fixed
TPP, meaning repeat budgets tuned at small scale are a floor rather than a ceiling. - Across domains, the optimal repetition count is strongly negatively correlated with a domain's final validation loss — domains the model already fits well can absorb more passes, while high-loss domains saturate sooner.
- The amount of unique domain data is only weakly related to the optimal repetition count, undercutting the common heuristic of setting epochs from corpus size alone.
- The headline practical claim — that repetition counts tuned on small proxy models at matched
TPPestimate the right setting for large models — is supported by trend direction rather than by reported absolute numbers or correlation coefficients here, and rests on validation loss rather than downstream benchmark quality, so how far the mild upward trend extrapolates beyond the tested size range remains open.
Forecast Collapse in Time-Series Foundation Models
Forecasting hourly returns for 1,000 US equities makes time-series models emit nearly flat predictions that rank stocks poorly, a failure the authors name forecast collapse and that mostly vanishes when the target is trading volume instead. Sweeping time-series foundation models (TSFMs), twelve deep forecasters, and 97 benchmark configurations traces the effect to target predictability plus per-series training objectives that never identify cross-series structure, exposing a calibration-versus-ranking tradeoff: squared error flattens forecasts, while optimizing cross-sectional correlation directly can inflate amplitude by over an order of magnitude. Their CalibRank objective balances the two and nearly triples cross-sectional correlation on Finance1K while keeping amplitude near the target, improving correlation on every model tested. The broader point is that per-series evaluation metrics can hide the cross-series failures that downstream decisions depend on.
Forecasting hourly returns for 1,000 US equities produces predictions that are nearly flat and rank stocks poorly by cross-sectional correlation — a failure the authors call forecast collapse — even though the same models forecasting trading volume in the same setting behave normally. The core claim is that collapse tracks target predictability, and that the standard squared-error objective is structurally unable to deliver both calibrated amplitude and usable cross-sectional ranking.
- The phenomenon is characterized across time-series foundation models, twelve deep-learning forecasters, and 97 public benchmark configurations, isolating two distinct causes: low predictability caps the amplitude any calibrated point forecast can take, while per-series training objectives never identify cross-series structure at all.
- This yields a calibration-ranking tradeoff — minimizing squared error drives forecasts toward flatness, whereas directly optimizing cross-sectional correlation recovers ranking but inflates forecast amplitude by more than an order of magnitude.
CalibRank, the proposed objective, mixes the two terms to balance calibration against ranking, and onFinance1Kit nearly triples cross-sectional correlation while holding forecast amplitude close to the target, improving correlation on every model tested.- The volume-versus-returns contrast is the cleanest evidence that this is about signal-to-noise rather than architecture, since identical models and data pipelines collapse on one target and not the other.
- The broader methodological point is that conventional per-series metrics such as MSE or MAE can look healthy while the cross-series structure that downstream ranking decisions actually depend on has already degenerated.
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Mobius-v0 splits a transformer's usual entanglement of knowledge and computation into a globally shared Memory built from feed-forward layers that stores knowledge vectors, and multiple Reasoners built from self-attention that repeatedly query that memory using hidden states as both cache and carrier. Because knowledge lives in one shared store while reasoning is applied iteratively, the architecture compresses knowledge more tightly and reuses reasoning operators. A 7B model trained from scratch matched a 7B transformer baseline's downstream scores using 62.6% of the baseline's training data, and Intern-S2-Mobius, continually pretrained from Qwen3.5-35B, matched downstream quality while delivering nearly 4x end-to-end inference speedup.
Mobius-v0 breaks the Transformer's layer-wise binding between feed-forward knowledge storage and attention-based reasoning, replacing per-layer FFNs with a single globally shared memory that every attention "reasoner" can query. The claim is that this decoupling buys both better knowledge compression and a shorter path to an answer, since latents can be refined against the full knowledge pool in a few layers rather than traversing every layer per token.
- Sharing one oversized key-value memory across layers gives the model an implicit backward residual connection — deep reasoners can reach shallow-layer knowledge, not just the reverse — plus dynamic latent reasoning in which hidden states iterate against the repository and decode several tokens at once; at scale the shared FFN is partitioned
MoE-style with sparse activation to keep the cost tolerable. - Trained from scratch as a 7B-A1B MoE on 1TB tokens,
Mobiusreaches the Transformer baseline's finalMMLUscore using only 62.6% of the data, which the authors attribute to less cross-layer redundancy in how knowledge is stored. Intern-S2-Mobius-35B, continually pre-trained fromQwen3.5-35B-A3Bon 1TB tokens plus SFT and RL, averages 67.88 vs 65.05 on general benchmarks (AIME 202695.31,HMMT 202685.51) and 52.14 vs 18.20 on scientific ones likeMol-InstructionsandMolecularIQ, while delivering roughly 4× end-to-end inference speedup.- The speedup comes almost entirely from shorter chains of thought rather than faster forward passes — on an
MMLU-Prolinear algebra item the model reaches the same correct answer in 516 tokens against 2,364, skipping the baseline's 1,147 tokens of repeated derivation and checking. - Weaknesses are real and acknowledged: the converted model loses to its own base on
UGD hard(73.02 vs 78.02) andHLE(19.11 vs 22.40), the mechanisms behind both the data efficiency and the shorter reasoning traces are hypotheses rather than established results, expert routing still carries a block-diagonal prior inherited from the source checkpoint, and the compositional-generalization evidence in the appendix comes from an unreleased newer architecture.
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Modern video generators can produce convincing footage of wars, disasters, and other emergencies, but existing detection benchmarks say little about how detectors behave on this kind of content or after it spreads online. RA-Bench pairs 1,830 real videos across 10 social-risk categories with 16,056 clips generated from four open-source and five closed-source generators, then evaluates seven traditional detectors, ten zero-shot multimodal models, and two multimodal large language models fine-tuned for the task. None of the three detector families generalizes consistently across the benchmark, generation quality and conditioning affect each family differently while source-level patterns stay stable across sampling seeds, and the videos that fool human judges are also the hardest for detectors — with social dissemination (re-encoding and sharing) making detection harder still.
Video generators can now fabricate convincing footage of wars, disasters, and public emergencies, but existing detection benchmarks are built on generic web video and say little about whether detectors hold up on socially consequential events. RA-Bench closes that gap by anchoring 1,830 real crisis clips across 10 social-risk categories and conditioning nine image-to-video generators on each clip's first frame, yielding 17,886 videos used to stress-test three detector families, generation properties, and survival through social sharing.
- Each real anchor is captioned with
Gemini-3.1-Pro-Preview, and that caption plus the anchor's first frame is fed to four open-source and five closed-source generators, so every fake shares scene semantics and a genuine opening frame with its real counterpart — mimicking the realistic attack where a true crisis photo seeds a fabricated video. - Detection collapses out of domain: the seven traditional detectors fall from published AUCs of
67.6–98.6%to source-level means of 43.9–57.3%, with 26 of 63 detector–generator pairs scoring below chance, and the public ranking barely transfers (Spearman 0.26), so no single detector leads across generators. - MLLM detectors fail differently —
Gemini-3.1-Pro-Previewtops the zero-shot models at only about63%BAcc while smallerQwen3.5variants flip from 19.7% to 100% fake-recall purely on prompt format, and fine-tunedSkyraturns out to key on a timestamp artifact: swapping absolute timestamps for frame indices over identical frames drops it from 68.5% to 54.4–54.9% BAcc, whileBusterX++calls almost everything real (FakeR4.1–9.1%). - The videos that fool people are the same ones that fool machines: 20 reviewers flagged 68.6% of open-source clips but only 52.9% of closed-source ones (
Seedance2.040.7%,Kling45.1%), and on the 633-clipRA-Bench-HumanProofsubset that all five reviewers called real, Gemini reaches just 54.7/54.5% BAcc and traditional detectors average 47.5% AUC — worse than a coin flip. - Real-world circulation is the hardest blow: the full
RA-Bench-LastMiledissemination simulation cuts mean fake-recall across five fine-tuned configurations from 46.0% to 1.4%, and more real-image conditioning (T2V → first-frame → first+last-frame I2V) steadily suppresses fine-tuned MLLM FakeR from 70.5% to 42.7% to 28.3%, though the study reports associations within one benchmark rather than causal effects and closed-source coverage is uneven because provider safety filters rejected some prompts.
Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Interactive game world models usually autoregress pixels or latents directly, which forces pose, geometry, and occlusion to be tracked implicitly and lets errors compound over long horizons. Marionette splits the job three ways: a two-stage autoregressive dynamics model predicts an explicit 276-dimensional 3D world state of articulated skeletons, metric root trajectories, and rotations; a zero-parameter graphics bridge turns that state into pose-control videos with world-space geometry and occlusion computed in closed form; and a control-conditioned video diffusion model paints photorealistic frames. Because behavior lives in the explicit state, it can be corrected there — left unconstrained, two generated characters drifted 21.2 m apart against roughly 5 m in recordings with a third of frames showing ground penetration, while adding just a terrain collider and a separation cap to the state cut penetration by 66% with no change to the observation model. Forcing a mismatched action stream shifted root-aligned joint error by 31% across 48 held-out segments, and routing appearance through the predicted state cost little visual fidelity (FVD 831 versus 799 for recorded poses).
Interactive game world models usually autoregress pixels or latents, leaving pose, geometry, and occlusion to be maintained implicitly by the same generative sequence, so errors compound over long rollouts. Marionette instead predicts an explicit 276-dimensional articulated world state, hands exact geometry to a zero-parameter renderer, and leaves only appearance to a video-diffusion model.
- Dynamics are split into a compact decision model (
ActionGPT, ~2.5M parameters, one discrete action token per frame per entity) and a larger animation model (PoseGPT, ~25M parameters, producing the 258-D body pose), after which a deterministic bridge integrates roots, places joints in metric world space, and rasterizes a three-channel geometry buffer thatWan2.2-Fun-5Brenders into photorealistic RGB. - Because the action is an explicit token, control is applied by overwriting it: feeding a temporally shuffled action stream degrades root-aligned joint error by 31% (0.272 m to 0.357 m) across 48 held-out segments, with the correct forced stream matching free-run quality at 0.272 m against 0.281 m.
- Routing every frame through the predicted state costs no measurable fidelity, with an FVD of 831 against 799 for recorded pose and 975 for an end-to-end pixel-autoregressive baseline trained on the same footage, though the bootstrap intervals for the baseline comparison overlap.
- Long-horizon failures live in the state and can be repaired there without touching the observation model: a free rollout drifts the two characters 21.2 m apart (recorded sessions stay near 5 m) with a third of frames showing ground penetration, while a terrain collider cuts the collision-frame ratio from 0.337 to 0.114 (66%) and a 6 m separation cap holds the pair at 5.1 m for 7% more foot-skate.
- The main limits are scope and drift: the dynamics model is trained on a single monster type's 173-action vocabulary while the appearance model spans 27 monsters, appearance identity decays across diffusion chunks since only a first frame anchors it, and a reported negative result shows a differentiable penetration penalty was gamed by stretching the skeleton (bone-length error rising from ~3% to ~13%), which is why constraints are imposed as inputs and post-hoc projections instead.
Applications 74
BCIJelly: An integrated ecosystem for brain-computer interface research
Brain-computer interface (BCI) research is slowed by incompatible data formats, one-off decoder implementations and hardware-specific deployment toolchains. BCIJelly bundles 18 curated datasets, 15 benchmark decoders, 80 reusable algorithm modules, an automated architecture search that builds task-specific decoders without manual design, and a toChip compiler targeting neuromorphic hardware, all in one Python framework with a graphical interface for non-programmers. The architecture search additionally runs in a closed-loop mode steered by a language model that reads task specifications, module descriptions and search history, and the stack is validated across five paradigms — motor, visual, speech, emotion and auditory — on recordings from humans, macaques and mice in single-task, multitask and cross-species settings.
Unknown Unknowns: Model Misspecification in Machine Learning for Physics
Inverse problems in particle physics and astronomy increasingly rely on models trained on simulation and deployed on real data, which raises the question of whether those models are wrong in unanticipated ways rather than merely whether they fit. The discussion frames misspecification as double-edged in physics — sometimes it is the discovery signal, sometimes a nuisance to absorb — and argues that a robust analysis absorbs the misspecifications one does not care about while preserving sensitivity to the ones one does. It surveys diagnostics for detecting misspecification and mitigation strategies, concluding that no single diagnostic can certify correct specification, so detection and mitigation must run as an iterative loop over a battery of complementary checks.
Building AI-Intensive Software with AI: Early Results and a Cautionary Tale on Measuring Development Cost
Reports on what AI-assisted software development actually costs are rare, and the few that exist are easy to get wrong in ways the final ratio never reveals. A six-person student team built a conversational onboarding assistant — retrieval-augmented code chat, guided tours, dependency graphs, technical-debt analysis — over one academic term with pervasive AI assistance, instrumented by a three-layer cost model tracking real AI spend, self-reported human effort, and a human-only counterfactual. The initially reported 19.4x cost advantage turned out to contain two independent mistakes, inferring per-token cost under a flat-rate subscription and pricing the counterfactual at the wrong regional labor rates, which together inflated the figure by roughly 2x; the corrected ratio is about 9.9x. The authors present the correction itself as the finding, since both errors are invisible in the headline number and plausibly common in similar reports.
Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
Work on clinical citations from chatbots has focused on fabricated references, leaving open whether the studies they do cite are the ones expert reviewers would pick. Three general-purpose assistants — Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5 — were prompted with clinical questions adapted from 20 Cochrane Database of Systematic Reviews questions under patient, clinician, and researcher personas, producing 720 responses benchmarked against each review's included and excluded study sets. Responses recovered 39.2% of the studies Cochrane reviewers included on average while citing only 5.0% of excluded ones, with recall varying sharply by model (63.1% for ChatGPT versus 17.3% for Gemini) and modestly by user role. After controlling for publication year, citation rate, and open-access status, trial sample size was the only independent predictor of retrieval (odds ratio 1.80 per log unit), indicating a systematic pull toward larger trials.
PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization
Macro placement drives a chip's final performance, power, and area, yet placers optimize half-perimeter wirelength, a proxy that recent benchmarking found nearly uncorrelated with post-route timing — all six AI placers evaluated on ChiPBench were worse than the hierarchical baseline. A label-fidelity study across ten circuits and four flow stages identifies post-global-routing metrics as the best trade-off between fidelity to final timing and label cost, so PPAPlace trains a dual-stream surrogate — graph attention over the netlist plus spatial convolution over the placement grid — on those labels and backpropagates predicted worst and total negative slack gradients to cell coordinates. Used both as a co-objective inside an analytical placer and as a projected-gradient refinement of macro positions, it delivers 22% better worst negative slack and 51% better total negative slack than the hierarchical baseline on five held-out circuits without retraining.
Recent Advances in Deep Learning-Based Drug-Target Binding Affinity Prediction
Drug-target binding affinity prediction is a core sub-problem of computational drug discovery where reported benchmark numbers have outrun real progress. Representative recent deep learning approaches are reviewed alongside seven widely used benchmark datasets and the standard evaluation metrics, with attention to architecture and molecular representation choices. The analysis finds that strong headline scores often reflect dataset bias and narrow evaluation protocols, and that most methods degrade substantially in cold-start settings involving unseen drugs or targets, pointing toward better dataset design, standardized evaluation, and multimodal representations.
HI-MeshGraphNets: Efficient and Accurate Mesh-based Physics Learning with Hierarchical Multi-scale Graph Neural Networks
Graph neural network surrogates for mesh-based simulation pass messages one hop per layer, so capturing long-range interactions on high-fidelity meshes demands deep processors that are slow, memory-hungry, and prone to over-smoothing. HI-MGN replaces the flat processor with a hierarchical one that coarsens the graph using farthest-point sampling and Voronoi partitioning while preserving the original mesh topology, letting information travel far in few layers, and reconstructs fine-resolution features with a learned graph interpolation network. On three structural and fluid benchmarks it delivers better accuracy than both MeshGraphNets and the Bi-Stride Multi-Scale GNN while cutting training time and peak memory.
Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Developmen
Team-level evidence for how AI coding assistants have changed open-source development practice is still thin, so this study tracks seven engineering metrics across all merged pull requests in two fast-moving AI infrastructure repositories, vLLM (18,290 pull requests) and SGLang (14,938), segmented into four eras. Throughput rose 21-fold and 17.9-fold respectively, and bot-authored pull requests accounted for under 0.2% of that growth, meaning the surge was overwhelmingly human-authored. Median cycle times in the latest era fell to roughly a day or less (with 90th-percentile tails above two weeks), comment density rose about four-fold with bots contributing an estimated 15-20% of that rise, contributor counts grew steadily, and pull request size stayed flat.
CipherSight: Robust Website Fingerprinting via Record-Resource Semantic Supervision under Distribution Shifts
Website fingerprinting infers which site a user visited from metadata leaking out of encrypted HTTPS traffic, but deployed classifiers degrade badly when traffic shifts across time and geography or when unseen sites appear. CipherSight learns representations from TLS records rather than raw TCP packet sequences, jointly encoding several record-level attributes in a hierarchical architecture that models both dependencies within a flow and interactions across concurrent flows, trained with a masked record modeling objective plus record-to-resource annotations used as privileged supervision. It reaches 95.41% accuracy over more than 2,000 website classes in the closed-world setting and stays above 90% accuracy under both temporal and geographic distribution drift, outperforming all evaluated baselines.
Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact
Text-to-SQL systems built on language models fail in a particularly dangerous way for production use: a hallucinated column or mis-aggregated total returns a fluent wrong number that looks identical to a right one, especially when the consumer is a dashboard or a tool-using agent that never inspects the query. The proposed architecture pairs a generative shell, which interprets underspecified input and phrases replies, with a deterministic kernel that matches fully specified questions against a bounded set of answerable question shapes and compiles them to queries, meeting at a user-read confirmation before any value is computed. The governing invariant is that a component able to fabricate may influence which question gets answered but never which value gets returned, so unanswerable requests are declined because they are structurally unrepresentable rather than filtered by a confidence estimate; the pattern is worked across three domains and backed by a two-year production case study compared against a fine-tuned parser and a tool-retrieval agent.
Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy
Short-video recommenders optimized for engagement systematically favor content that grabs attention within seconds, and there is growing evidence that heavy exposure to such shallow content affects cognitive engagement and well-being. The proposed Content Depth Score rates how far a video is expected to stimulate higher-order cognitive processes on a seven-level scale grounded in cognitive psychology and learning theory, and SCOPE-Bench applies it as annotations over 150K videos from a large open-source dataset. Evaluating 13 representative recommenders reveals a consistent pull toward shallow content, and algorithms that do surface cognitively deep videos are only marginally better than random selection.
Buy the Rumor, Sell the News: When Is News Priced In?
Two market adages hold that news is already priced in by publication and that traders buy the rumor and sell the news, both placing the price move before rather than after publication. The test covers 4.57 million financial news articles on roughly 3,000 US stocks from 2023 to 2026, with a large language model teacher distilled by active learning into a compact classifier that assigns 17 event tags and five attributes, articles clustered into stories to distinguish first reports from follow-ups, and beta-adjusted abnormal returns measured around 1.68 million stock-day events against 364,405 neutral-sentiment placebo events. The cumulative move in the news direction by the close of publication day is 2.8 times its value 20 days later, and for rumor-flagged events the rumor day captures the entire move while confirmation adds nothing; quantified fundamentals such as earnings and guidance keep drifting for weeks while soft story-driven news gives its move back, and publicity raises volatility before publication and lowers it after.
HERMES: a multi-agent framework for structured knowledge extraction from ultra-long documents in geoscience
Much authoritative geoscience knowledge sits in decades-old monographs whose unstructured prose and complex layouts block computational reuse. HERMES is a multi-agent extraction framework in which a coordinating language model applies domain constraints, validation rules, and evidence tracing over parsed text, tables, figures, and captions of ultra-long documents. Applied to the 55-volume Treatise on Invertebrate Paleontology, it produced a public database of 32,277 fossil taxonomic entities and 451,878 attributes at roughly 0.90 F1 for entities and 0.91 for attributes, about six times faster per volume than the manual baseline, and transferred without retraining to palaeomagnetism and geochemistry documents.
Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency
Studies of language-model-based automated program repair mostly report aggregate fix rates, leaving open how performance depends on how complicated a bug is, how precisely the fault was localized, how much reasoning effort is spent, and what it all costs. Two repair techniques, ChatRepair and CodeCorrector, are run across DeepSeek, GPT, and Llama models under varying bug complexity and localization precision with statistical analysis of the results. Over half of moderately complex bugs are repaired by low-cost models, imprecise fault localization widens the gap between repair techniques substantially, and spending more — pricier models or heavier reasoning settings — does not reliably improve cost-efficiency: GPT-5 fixes 7 and 39 more complex bugs than DeepSeek-V4-pro and DeepSeek-V3.2 respectively, while DeepSeek-V3.2 is the most cost-efficient overall.
When Denoising Hurts: Rethinking the Terminal Step of Diffusion Time Series Forecasters -- Extended Version
Diffusion forecasters are usually run to the end of the reverse process on the assumption that every denoising step refines the prediction. Tracking forecast quality across the reverse trajectory shows the opposite at the tail: broad temporal structure is recovered while noise is still high, and continued low-noise refinement introduces statistical drift that degrades accuracy — which also explains why prior work gravitated to narrow architecture and schedule choices. The proposed fix is a label-free global stopping criterion that detects when to terminate, plus a Bernoulli timestep sampler that concentrates training on the high-noise region where inference now ends while still covering the full schedule; experiments on eight real-world datasets show both faster inference and better accuracy.
Forecast Collapse in Time-Series Foundation Models
Forecasting hourly returns for 1,000 US equities makes time-series models emit nearly flat predictions that rank stocks poorly, a failure the authors name forecast collapse and that mostly vanishes when the target is trading volume instead. Sweeping time-series foundation models (TSFMs), twelve deep forecasters, and 97 benchmark configurations traces the effect to target predictability plus per-series training objectives that never identify cross-series structure, exposing a calibration-versus-ranking tradeoff: squared error flattens forecasts, while optimizing cross-sectional correlation directly can inflate amplitude by over an order of magnitude. Their CalibRank objective balances the two and nearly triples cross-sectional correlation on Finance1K while keeping amplitude near the target, improving correlation on every model tested. The broader point is that per-series evaluation metrics can hide the cross-series failures that downstream decisions depend on.
From Fixed Grids to Moving Particles:A Transferable Latent Operator for Fluid Dynamics
Fluid simulation data is usually available as Eulerian fields on fixed grids, while Lagrangian particle trajectories — needed to describe transport — are scarce, so neural operators are almost always trained and tested in the Eulerian view. The Transferable Latent Operator (TLO) separates latent flow evolution from coordinate-dependent decoding: querying the evolving latent state at grid points yields Eulerian fields, while querying velocities at particle positions and integrating them forward produces Lagrangian rollouts from the same model. Across five fluid-dynamics benchmarks TLO beats existing neural operators on Eulerian prediction and generalizes zero-shot to Lagrangian particle rollout without any Lagrangian supervision, improving further with a small amount of Lagrangian fine-tuning.
BGA: A noise-immune neural distillation framework for malicious signature extraction in high-entropy encrypted flows
Encrypted TLS 1.3 traffic is high-entropy enough that attention-based detectors spread their weight across cryptographic noise instead of attack signal. BGA first uses analysis of variance to separate discriminative control-plane features such as industrial setpoints from that noise, then trains a Wasserstein GAN with gradient penalty to synthesize rare-class flows within an 86,878-record corpus, and finally combines a bidirectional LSTM with an adaptive gated multi-head attention layer that acts as a filter on encryption artifacts. Recall on rare Malicious State Command Injection attacks rises by 43.2%, all key metrics exceed 95.2% on the CIC-IDS-2018 and Edge-IIoT benchmarks, and noise-injection tests give an 8.57% robustness margin over a vanilla Transformer at 0.282 ms inference latency, which the authors extrapolate to roughly 1.7 ms on ARM edge gateways.
MINT: A Universal Zero-Shot Predictor for Transaction Data
Banks turn transaction sequences into embeddings with payments foundation models, but those embeddings feed task-specific heads and cannot answer novel questions zero-shot, whereas LLM-based approaches lose predictive signal and pay heavily for serializing transactions into text. The Multimodal Instruction Network for Transactions (MINT) wires a pretrained transaction sequence encoder into a decoder-only LLM through lightweight embedding injection, transaction-language alignment, and instruction tuning. It reports state-of-the-art predictive question-answering both in-distribution and out-of-distribution while substantially cutting input tokens, latency, and memory relative to text-serialization baselines, and analyses of representations, alignment strategies, training data, and history length support compact embeddings over serialized text for this kind of multimodal reasoning.
How Much Do Legal RAG Systems Still Hallucinate?
Ungrounded answers carry real consequences in law, so this study audits hallucination in eight legal retrieval-augmented generation (RAG) systems over two corpora: the GDPR in English and a national civil law in French. Evaluation runs at both claim and answer level, reporting hallucination density and severity broken down by question category and user persona, with findings validated on 142 questions written by legal experts. Hallucination remains pervasive, affecting under 10% of responses for the best systems but close to half for the worst, and false-premise questions — those built on incorrect assumptions the system should reject — trigger especially high rates on the expert-authored set.
Non-Parametric Spatiotemporal Trajectory Prediction via State-Conditioned Transition Sampling
Multi-modal trajectory prediction is normally handled by trained sequence models, which need substantial historical data and GPU training before they can serve a new geographic region. The alternative here is entirely training-free: a lookup table of historical state-to-next-position transitions, queried with a product kernel over spatial proximity, bearing, speed, and temporal context, with two inference modes over the same table — diversity-penalized sampling for covering distinct plausible routes, and beam search for the single most likely path. On the TrAISformer benchmark of Danish maritime AIS data it matches a 57M-parameter transformer at full data availability with zero learned parameters and no GPU, and stays stable down to 10% of the training data where the transformer degrades catastrophically, which points to deployment in new regions from an order of magnitude less history.
Catching the Imposter: Self-Supervised Learning of Physical Coherence with Cross-Entity Feature Permutations
Scientific measurements describe entities whose variables are coupled by physical law, a structure that standard self-supervised objectives leave unused. The imposter pretext task swaps a subset of one entity's feature values with genuine observations donated by a different entity and trains the encoder to spot which features were swapped; since every donated value is individually plausible, the only way to solve it is to learn cross-feature physical dependencies. Evaluated on 21 variables from global ERA5-Land reanalysis across seven downstream tasks in climate classification, carbon flux estimation, and streamflow prediction — the first systematic comparison of self-supervised objectives for land-surface modelling under a shared architecture and pre-training budget — the results show no single objective dominates, with the best pretext task depending on the downstream family and imposter contributing information complementary to existing objectives.
Universal Thermodynamic Interatomic Potentials for Crystalline Materials
Solid-state phase stability is governed by free energies, but high-throughput materials screening still leans on ground-state energies because free energies require ensemble averages that are expensive to compute. The thermodynamic interatomic potential (TIP) extends a standard interatomic potential from static energy to a thermodynamically consistent Gibbs free energy model, so responses to temperature and pressure follow by automatic differentiation; the implementation TIP[UMA] builds on the universal potential UMA, trains on free energies spanning quasi-harmonic to molecular-dynamics fidelity, and is calibrated against higher-resolution calculations or experiment. A single evaluation returns a crystal's equation of state and locates phase transitions among competing branches, including dynamically stabilized phases, and fine-tuning extends coverage to alloy solubility limits and miscibility gaps.
51 more specialized papers
- L-FNO: Lorentzian Fourier Neural Operator for Stochastic Event Dynamics Songhee Kang, Jihoon Kang
- Proxy-Validated LLM UX Micro-Simulations: An Artifact-First Protocol for Early-Stage Decision Support Alexandre Cristov\~ao Maiorano
- Interactive Analysis of Global Explanations using Aggregated Class Activation Maps for Network Data Igor Cherepanov, David Sessler, Alex Ulmer et al.
- From Prediction to Intervention: Personalized Meal-Level Glucose Regulation via an LLM Agent Mingyu Huang, Weiqing Min, Ying Jin et al.
- Context Aware AI Assistant and AR Interface for Lunar Extravehicular Activity (EVA) Procedural Guidance Rodrigo Gallardo, Qilmeg Doudatcz, Ganit Goldstein et al.
- How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights Himanshu Tripathi, Kaushik Roy, Subash Neupane et al.
- Learning Unsteady Aneurysm Hemodynamics with Physics-Informed DeepONets Oscar L. Cruz-Gonzalez, Val\'erie Deplano, Badih Ghattas
- Contrastive Learning for Interpretable Anomaly Detection at Collider Experiments Haoyi Jia, Sagar Addepalli, Julia Gonski
- EEG-PRISM: Physiologically-Grounded Interpretability of Predictions by EEG Foundation Models Deeksha M Shama, Punnisa Amornsirikul, Archana Venkataraman
- TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials Fatema Tuj Johora Faria, Mukaffi Bin Moin, M. F. Mridha et al.
- Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation Divya Vetticaden, Arya Gupta, Julian Nyarko et al.
- BM25-Augmented Many-Shot Translation for Low-Resource North-Eastern Indian Languages Aashish Dhawan, Christopher Driggers-Ellis, Dzmitry Kasinets et al.
- Architecture and Affordances of PLAUD: Performative Latents and Unsupervised DDSP B{\l}a\.zej Kotowski, Frederic Font
- Data-driven techniques for translational neuroscience and personalized neuro-health Vishal Subedi, Shashipraba N. K. Rajakaruna, Pratyusha Sarkar et al.
- Stochastic Control Policies for Robust Molecular Transition Path Sampling Jingqian Liu, Yu-Hsiang Wang, Yanru Qu et al.
- Optimal Power Allocation and AI Receiver Design for Superimposed DMRS and Data Transmission Sha Hu, Zhongwang Fu
- SPEAR: Structure Property Explainability with Attention Regularization Aditya Raghavan, Utkarsh Pratiush, Dalton A. Pearl et al.
- Fashion Outfit Generation via Unified Sequential Composition Models Kaicheng Pang, Xingxing Zou, Ruohan Xu et al.
- MedMix: Specialization-Consistent Federated Sparse MoEs under Modality Heterogeneity Adiba Orzikulova, Dong Min Kim, Jaehong Yoon et al.
- Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning Chun-Hua Lin, Samuel Yen-Chi Chen, Yu-Chao Hsu et al.
- Probabilistic indirect models for undrained shear strength: addressing significant data missing and variability with advanced imputation and machine learning techniques Haibin Xiong, Shaoheng Dai, Peng Lan et al.
- Deep Vision in Smart Manufacturing: MODERN Framework for Intelligent Quality Monitoring and Diagnosis Yicheng Kang, Yuling Jiao, Xin Geng et al.
- CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification Bingxin Yu, Xueli Wang, Jerry Zhou et al.
- Musical Mirrors: The LLM as Sounding Board in Songwriting Xiao Xiao
- Simulation-Driven Vehicular Traffic Data Augmentation: Extending Sensor Coverage Through Virtual Sensing Davide Andrea Guastella, Eladio Montero Porras, Evangelos Pournaras et al.
- EchoRec: Multi-Item Prediction-Empowered Generative Recommendation via Cycle-Consistent Preference Alignment Haokai Ma, Aoqi Hu, Yueao Xing et al.
- Residual Dominance as a Structural Account of Last-Item Reliance in Causal Self-Attention Recommenders Keito Kozaki, Keigo Sakurai, Ren Togo et al.
- Model-agnostic Retrieval-Augmented Extended Forecasting for time series Juan Pablo Villa Serna, Rohan Asthana, Vasileios Belagiannis
- Voxel-based 3D Facies Segmentation from Seismic Data: A Comparative Study Duc-Thanh Pham, Minh-Tan Pham, Anh Nguyen et al.
- Benchmarking data-driven material models on the classic Treloar dataset Hagen Holthusen, Moritz Flaschel, Denisa Martonov\'a et al.
- Reaction-Transformation-Aware Flow Matching for Generalizable Transition State Generation Kaipeng Zeng, Wenxi Zhai, Shengrui Xu et al.
- Fixed-Budget Gaussian Volume Encoding with Structure-Aware Allocation Michael R. Martin, Joseph Insley, Victor A. Mateevitsi et al.
- Smart routes: a system for development and comparison of algorithms for solving vehicle routing problems with realistic constraints Andrew Soroka, German Mikhelson, Alexander Mescheryakov et al.
- Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach Sheng Hong, Xuanqi Wang, Jiacheng Wang et al.
- Structure-Guided Spatiotemporal Attention Graph Neural Network for Traffic Flow Prediction Xuanmian He, Can Li, Wanjing Ma
- Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response Timothy C. Pearce, David J. T. Smith, Alec Dobney et al.
- Pairton: Iterative Reconstruction of Short-Lived Particles Andreas Hermansen, Chris Scheulen, Tobias Golling
- Acoustic UAV Detection in Battlefield Scenarios: Handling Noise, Domain Shift, and Weak Labels Vadym Vilhurin, Volodymyr Sydorskyi, Andrii Shevtsov
- Conditional Neural Optimal Transport for Predicting Cellular Phenotypes from Molecular Structure Gauthier Avit\'e, Maxime Sanchez-Renauld, Nicolas Bourriez et al.
- Intelligent Detection of Mechanical, Electrical, and Plumbing (MEP) Metrics Based on 2D Floor Plans Tarandeep Singh Mandhiratta, ANK Zaman, Abdul-Rahman Mawlood-Yunis
- Program-space Diffusion for Morphology-to-Transcriptomics Prediction Ruyter Swann, Dorent Reuben, Racoceanu Daniel
- Mind the Long Tail: Understanding the Difficulty of Delay Detection in Business Processes Keyvan Amiri Elyasi, Lukas Kirchdorfer, Heiner Stuckenschmidt
- A Hybrid LLM-Based Framework for Automated Security Annotation Generation in Business Process Models Md Kamrul Islam, Tiphaine Henry, Mattia Salnitri et al.
- A Survey of Large Models in Sports Yichen Xu, Jianzhe Ma, Chuhan Wang et al.
- From Style Replication to Style Exploration: Enabling Art Style Exploration with Analyze-Experiment-Resituate Framework Wen-Fan Wang, TsaiHsuan Lin, Chi-Lan Yang et al.
- CytoBERT: A Foundation Model for Cytometry Data Syed Abdul Haseeb Qadri, Bjarne C. Hiller, Felix Blanke et al.
- Shift Aware Transfer Learning with Adaptive Dual-Encoder Fusion for PM Forecasting in Data-Limited Environments Shahab Band, Hamed Mohammadi
- Optimal Scheduling of Road Maintenance Jobs Considering Impact on Traffic Flows Charitha Nandepu, Lohitha Kalepu, Gabriele Ciavarella et al.
- Generating Benchmark Health Data Using a Tabular Diffusion Transformer Hao Yan, Lisa Pilgram, Dan Liu et al.
- Learning-to-Transition for Large-scale and High-Order MIMO Detection Yubo Zhang, Yiyao Liu, Xiaodong Wang
- Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils Karel Becerra, Boris Mederos, Dean Snow et al.
Large Language Models 45
Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking
Which layers of a sparse Mixture-of-Experts (MoE) model can be pruned is poorly characterized, so the authors mask low-magnitude experts layer by layer in Qwen3.6-35B-A3B (40 MoE layers, 256 experts each, top-8 routing) and score output quality on the XLCoST cross-lingual code-translation benchmark at 100-, 300- and 500-prompt scales. Sensitivity proves strongly depth-dependent: early and middle layers degrade badly under masking, while the last five layers tolerate aggressive removal of low-magnitude experts. Masking 640 of 10,240 experts confined to layers 35-39 retains 419/500 good-or-similar outputs, against 150/300 for flat 30% masking across all layers. Narrowing routing from top-8 to top-6 active experts cut wall-clock time with no quality loss on a small probe, but did not compose cleanly with heavy masking.
Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required
Post-training reports, model cards and blog posts routinely treat SWE-bench and LiveCodeBench scores as evidence of general coding ability, and the authors build a Django-based multi-task benchmark suite to test whether that inference survives contact with different tasks. Evaluating foundation models alongside checkpoints post-trained on SWE-bench trajectories, they find benchmark rankings frequently fail to generalize, with SWE-bench optimization yielding limited or no gains on the Django tasks or on LiveCodeBench, and fine-tuning on individual Django task types likewise failing to transfer. Their recommendation is differentiated evaluation — holistic assessment for frontier models, multi-task suites for research, human-in-the-loop studies for narrow applications — plus a capability taxonomy and sustained benchmark maintenance instead of one-off releases.
Modular Cognitive Architecture Emerges in Large Language Models
Human brains show pronounced functional specialization, with separate networks handling language, formal reasoning, reasoning about other minds, and reasoning about the physical world; the question here is whether that modularity is a necessary property of intelligent systems or an accident of biology. Circuit analyses across 46 tasks spanning those four cognitive domains trace which neurons each task recruits inside large language models. Tasks that share a brain network in humans recruit overlapping neurons in the models, while tasks drawing on different human networks recruit distinct ones, which the authors read as modular organization emerging convergently under a very different optimization process.
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
Studies of production language-model serving so far cover short windows and few models, leaving little visibility into how traffic evolves or how user-model interaction shapes it. The authors characterize a full year of production traffic from Chutes, spanning many models and users including long-tail ones, analyzed from aggregate, temporal, model-level and user-level perspectives with attention to caching and load-balancing behavior. The complete one-year trace will be released alongside the paper so serving-systems research can benchmark against real rather than sampled or synthetic workloads.
Jais 2: A Family of Arabic-Centric Open Large Language Models
Jais 2 is a family of Arabic-centric open models from MBZUAI, Cerebras and Inception, comprising what the authors report as the largest open Arabic-centric model trained from scratch at 70B parameters plus a competitive 8B variant. A custom Arabic-centric vocabulary together with an optimized architecture and training recipe let the models reach strong Arabic results on a substantially smaller token budget than comparable systems, leading the evaluated open models on OALL2 and AraGen while staying competitive in English and performing well on culturally grounded benchmarks covering poetry, religion, cuisine and dream interpretation. Weights are released on HuggingFace under a commercially permissive license, and the 70B chat deployment runs on Cerebras hardware at up to 2,000 tokens per second.
IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering
Multi-hop question answering buries retrieval-augmented systems in long, noisy contexts, and existing prompt compressors are designed for single-turn queries so they discard evidence that only becomes relevant at a later reasoning step. IterCOMP is training-free: it decomposes retrieved documents into evidence segments, judges whether the question is answerable yet, and generates targeted follow-up questions to pull in missing evidence over successive rounds, ending with a compact reasoning-oriented prompt. On MuSiQue, 2WikiMultiHopQA and HotpotQA it improves exact-match and F1 over existing compression baselines while cutting the token budget, with the advantage growing as reasoning complexity increases.
Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors
Confidently wrong answers from language models are usually read as a sign that the underlying inference is shaky, but this work tests whether some are instead locally stable — unchanged by small perturbations. Two diagnostics are combined: an output-level audit score that ranks domains by confidence variation and overconfident mistakes under a forced-answer baseline, and an internal probe measuring how far hidden states move. Self-critical prompting reliably reduced hidden-state sensitivity across layers in three open-weight models, yet overconfident errors were not measurably more locally sensitive than confidently correct answers, implying that prompting stabilizes representations without actually fixing calibration.
Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning
Transferring capability from a large donor model into a much smaller recipient normally requires aligning neurons or architectures, which is hard when the two differ in scale and structure. Activation-Prune-Merge (APM) instead builds task-conditioned activation maps on the donor, uses them to select the most salient layers, hidden dimensions, attention heads, and MLP neurons, prunes the donor down to the recipient's shape, and blends the resulting slice in with a very small interpolation coefficient — no training required. Across 16 benchmarks covering reasoning, math, code generation, instruction following, and classification, average accuracy of a 3B recipient rose from 55.5% to 60.6%, with RTE climbing from 64.3% to 82.3% and BoolQ from 70.8% to 79.2%.
No Universal Signal Predicts Sample-Level LLM Regression under Version Updates
A new model version usually beats the old one on aggregate scores while still breaking individual cases that previously worked, and the question here is whether such per-sample regressions can be predicted at inference time. Single-model signals (confidence, logit margin, attention entropy) are pitted against cross-version signals (output and token-level KL divergence, likelihood drift, representation drift) under a common test that isolates each signal's gain over a plain confidence baseline, spanning six benchmarks in multiple-choice question answering, math reasoning, and code generation across six model update pairs. No signal won universally — confidence dominated on multiple-choice and easy math, while likelihood and KL divergence signals helped most on harder math and code — but some cross-version signals stayed informative without labels where confidence failed, enabling a proof-of-concept fallback that routes high-risk inputs back to the previous model.
The Query Knows What to Forget: A Second Erase Direction for Linear Attention
Linear attention compresses history into a fixed-size state, so at long context many stored items crowd each other and retrieval degrades. Delta-rule models including Gated DeltaNet-2 derive their erase vector from the current token's key, but read interference is measured through the query, a direction the key-based erase cannot touch. The Query-derived Erase Direction (QED) adds a second erase component taken from the query and made orthogonal to the key, using the part of the state a key-directed edit cannot reach to cancel stale content along the query; it improves retrieval at every length beyond the training window and roughly doubles usable context length on S-NIAH-1.
From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models
A retrospective covering October 2018 to July 2026 tracks how language models went from BERT to agents that solve competition math and write software, reporting that the ability to resolve real coding issues improved nearly sixfold per year since late 2024 while cost per unit of capability collapsed, with budget-tier models matching flagship capability at one to six dollars per million tokens. It also documents fragmentation of the frontier into task-targeted models, with different systems leading on frontend coding, repository-level coding, and terminal tasks. A companion grade-school math experiment on Qwen2.5 shows basic decoding solving 58 of 100 problems against up to 79 with advanced sampling, and a confidence ranker placing 47 correct answers in its top 50, with all materials released publicly.
Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT
Translating legacy C into safe, idiomatic Rust would remove whole classes of memory-safety bugs, but off-the-shelf language models are weak at it because general pretraining emphasizes neither idiomatic Rust nor cross-language semantic equivalence nor repairing compiler feedback. The reported recipe applies a three-stage curriculum to Qwen3-27B: continued pretraining on Rust-heavy corpora, supervised fine-tuning on microsoft/Verus_Training_Data to instill debugging and self-repair, and task-specific fine-tuning on paired C/Rust LeetCode solutions. Evaluation runs through SACTOR, an agentic static-analysis-guided framework that does two-phase unidiomatic-to-idiomatic translation with foreign-function-interface end-to-end tests, reporting success rate, Clippy lint counts and unsafe-code fraction as idiomaticity measures, and failure-mode breakdowns against baseline Qwen3-27B and other models.
Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline
Non-functional requirements in code-generation prompts are usually a terse one-line phrase, and it is unclear whether spelling them out against a quality standard helps. Four requirements — performance, error handling, code smell, readability — were expressed three ways (a one-line baseline, rich prose grounded in the ISO/IEC 25010 quality model, and structured JSON with the same ISO content) across ten prompt variations each, evaluated on HumanEval and HumanEval-ET under a fixed model snapshot with paired non-parametric tests. ISO grounding improved static quality proxies and reduced sensitivity to prompt wording but did not reliably improve functional correctness, and for error handling the extended-test pass rate actually fell, suggesting defensive coding conflicts with exact-output benchmarks. Holding ISO content constant, prose and JSON differed negligibly in correctness, so the semantic content matters more than the serialization format.
The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference
Two GPU kernels implementing the same scaled INT8 general matrix multiply interface are normally assumed interchangeable, so swapping only the INT8 linear kernel (CUTLASS versus Triton) inside vLLM — with checkpoint, prompts, hardware, decoding and quantization config held fixed — should not change output. Each arm reproduced itself bit-for-bit across restarts, yet the two arms agreed on no generated sequence in any end-to-end comparison run (0 of 8, 0 of 16, 0 of 64), even though the INT32 accumulator is provably exact and order-independent under a verified no-overflow bound. Feeding both kernels identical operands from every linear layer of Qwen3-1.7B and 8B produced bit-identical outputs under power-of-two scales and differences of at most one bfloat16 spacing under the real scales, localizing the divergence to scale application and output rounding after the accumulator; applying that as a probe restored full bitwise agreement end to end. Cross-implementation FP8 shows a different signature that grows with reduction depth, and teacher-forced replay shows flips concentrate at small logit margins, predicting flip risk with an ROC-AUC of 0.94.
Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future Directions
Federated learning lets clients train together without pooling raw data, and combining it with prompt-based adaptation offers a way to specialize large language models without the compute or data centralization that full federated fine-tuning requires. This survey organizes federated prompt learning across the model lifecycle — pre-training, fine-tuning, and deployment — covering how it differs from conventional federated learning and full-model federated fine-tuning, and cataloging defense mechanisms for the attacks it inherits. The comparison centers on the trade-offs each approach makes among accuracy, communication cost, computational overhead, scalability, personalization, and client heterogeneity, closing with open security, privacy, robustness, and systems challenges.
Bootstrapping Niche Multilingual Code Translation via Reinforcement Learning with Execution-Based Verifiable Supervision
Neural code translation has concentrated on a few popular languages such as C++, Java, and Python, leaving many-to-many translation among less common languages short of parallel supervision and prone to producing plausible but non-executable code. The pipeline grows verified seed Python programs into an execution-validated multilingual pool, labels candidate translations by whether they actually run, trains a reward model on the resulting preferences, and optimizes base models with GRPO over 600 directed language pairs. A new execution-based benchmark, HumanEval-X++, extends HumanEval-X to this wider language space, and the 4B Qwen-3.5 model improves by 13% on average across all languages, with a 21% gain on mid-tier languages.
Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification
Synthetic training text from large language models varies widely in usefulness: some generated samples land inside the correct class region of an embedding space while others drift to the periphery or into other classes. The proposed filter scores each generated sample by its Euclidean distance to real class examples in a sentence embedding space and converts those scores into soft sample weights for classifier training. Across 13 datasets, 5 classifiers, 10 augmentation methods, and more than 6,700 configurations, the method beats SMOTE by 2.61 percentage points with an 88.9% win rate and transfers unchanged to named entity recognition for a 9.26-point gain; notably, the simplest distance-based filter consistently outperforms more elaborate multi-criteria variants.
CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing
Diffusion language models generate faster by filling several masked positions per forward pass, but aggressive parallelism makes their early denoising predictions unreliable and those errors compound downstream. Consistency Forcing is a distillation method that trains on pre-collected self-rollout trajectories so early-stage mask predictions align with later-stage ones, using a confidence-adaptive Kullback-Leibler objective that blends the strengths of forward and reverse divergence, with theory explaining why the consistency target approximately minimizes early-stage prediction error. The same formulation covers both mask-to-token and edit-capable decoding, and experiments on LLaDA variants show improved speed-quality trade-offs that widen under high-parallelism decoding budgets.
Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion
Creative-writing training data for large language models is dominated by stories, so models trained on it struggle with the structural and formatting conventions of other creative formats. The proposed framework separates thematic breadth from genre control by pairing human-authored story prompts as creative seeds with manually curated genre attributes that enforce distinct structure and style, then prompts strong models for query-response pairs and quality-filters them, producing the Multi-Genre Collection of 50K examples across 13 genres including rap, lyrics, scripts, game design, and character design. Models fine-tuned on this corpus beat base models, writing-specialized baselines, and models trained on existing writing corpora on out-of-distribution benchmarks, and genre-count ablations indicate that controlled genre expansion, not more story data, is the main driver of the gains.
Repair, Not Improvement: Decomposing Constrained Decoding in Tool-Call Abstention
Prior accounting of constrained generation found the decoder contributes little for format constraints but explicitly warned against extrapolating to cases where a constraint encodes correctness, naming function calling as one; this study measures that excluded case, focusing on when a model should decline to call any tool. Three conditions over one byte-identical prompt separate a grammar's two jobs — fixing where generation stops and which tokens may be emitted — across open-weight models from 0.6B to 4B on matched English and Korean items. Against an unconstrained decoder, constrained decoding is negative on abstention in four of six cells with intervals excluding zero, worst -29.5 points, and positive in none, and the total is a sum of opposing effects: on the smallest model in Korean the stop token costs -20.0 while the enum returns +19.5. What the grammar recovers is form rather than judgment, since 545 of 698 repaired abstentions had no readable answer to begin with, and both preregistered language claims fail.
QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction
Quantization-aware training (QAT) computes losses and surrogate gradients from a lossy reconstruction of latent full-precision weights while updating the latent weights themselves, a mismatch that raises the final loss floor; second-order post-training methods close a similar gap but take hours per pass and cannot be repeated as weights evolve. QUASAR performs lightweight loss-aware reconstruction inside the training loop, using an exponential moving average of squared gradients as online saliency estimates, searching a small set of clipping ranges, and fitting affine dequantizers by saliency-weighted least squares; the accompanying analysis shows loss-aware reconstruction error is the only reconstruction-dependent term in the QAT convergence bound. It changes only training, keeps standard deployment formats including integer quantization and NVFP4 with no inference overhead, and on Qwen3 and Llama-3.1 cuts held-out KL divergence by at least 10% at 3 and 4 bits and 29% at 2 bits, improving average accuracy across eight tasks by 3.5 to 4.3 percentage points at 2 bits.
Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead
Nanbeige4.2-3B is a 3B-parameter agentic model whose Looped Transformer reuses one layer stack for a second forward pass, buying effective depth without extra parameters but roughly doubling peak attention memory. Running it on Apple Silicon surfaced five independent bugs blocking the released checkpoint under Hugging Face transformers, including a silently zeroed rotary position embedding buffer and calls to removed cache APIs; a chunked-prefill strategy then addresses the memory penalty, extending allowable context width by 2.7x on 32 GiB of shared memory. After further system-prompt and Metal-backend memory patches, the debugged model completes up to 30% of real agentic tasks on a subset of MCPMark (up from 0%) and is near-perfect on single tool calls in BFCL while failing most multi-tool tests; the patched checkpoint and harnesses are released.
Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model
Large reasoning models generate long chains of thought that are expensive to serve, and production throughput requires batching, but existing training-free adaptive pruning methods collapse in that regime: a batch shares one pruning mask, so activations are aggregated across samples while the pruning threshold was calibrated offline on unaggregated activations, causing the realized sparsity to drift. The proposed method replaces threshold selection with periodic top-k selection over aggregated importance scores, which is immune to the distribution shift aggregation induces and runs once per update period rather than per token, and adds an activation memory that accumulates importance across phases because important neurons re-fire periodically during long generations. On DeepSeek-R1-Distill-Qwen-7B at batch size 4 and 50% target sparsity it beats the previous state-of-the-art adaptive pruning method by 39.7 percentage points in average accuracy, reaching a 1.40x speedup over dense inference at 50% actual sparsity.
Scaling Domain Data Repetition in LLM Pretraining
As models grow, token budgets grow with them to keep a sensible tokens-per-parameter ratio, but scarce high-quality domain data cannot scale the same way and gets diluted in the mixture; repeating it counteracts dilution at the risk of overfitting. Sweeping this trade-off under practical scaling, where token budget grows with model size, yields two findings: at fixed tokens-per-parameter the optimal repetition count mildly increases with model size, and across domains the optimal count correlates strongly and negatively with a domain's final validation loss, while the amount of unique domain data barely matters. The practical implication is that repetition counts tuned on small proxy models at the same tokens-per-parameter ratio carry over to larger runs.
The conditional superiority of fast silicon sampling
Silicon sampling — using language models to stand in for human survey respondents — sometimes reproduces population statistics well, and the question here is whether a cheaper, faster sampling mode sacrifices that fidelity. Fast and slow modes are compared against a nationally representative sample of Singaporean respondents. Both modes estimate population means moderately well but understate opinion variance and distort the latent contextual structure behind human views, so the method warrants caution overall; conditional on those limits, the fast mode is uniformly at least as faithful as the slow mode while using far less compute and wall-clock time.
P2Skill: Privacy Preserving Skill Distillation for Cloud-Local LLM Inference Systems
Splitting inference between a phone-resident small model and a cloud model only protects users if cloud-bound requests are stripped of personally identifiable information (PII), and existing masking or perturbation schemes distort meaning or demand extra training. P2Skill instead hands a local small language model (SLM) a set of prompt-defined skills — query decomposition, PII-aware routing, paraphrasing, and reconstruction — that a cloud LLM iteratively rewrites based on observed execution failures, so the system needs no privacy-specific fine-tuning and no learned PII detector. On a four-domain benchmark it reports 1.69× and 3.66× higher privacy-preserved inference quality than prior baselines.
Retrieval Grounding Latent Reasoning for Dense Retrieval
Reasoning-intensive retrieval needs embeddings that encode not just topical similarity but the inference required to judge relevance under an instruction, and models that bolt reasoning onto embeddings are typically trained end-to-end on the retrieval loss alone — leaving room for shortcut latent trajectories that score well without adding anything. RGLT (Retrieval Grounding Latent Reasoning) runs non-autoregressive reasoning in hidden space over instruction-conditioned silent tokens, shaping intermediate states with stage-wise chain-of-thought reconstruction distilled from explicit reasoning and assigning retrieval-effect credit so each latent step is optimized for its incremental retrieval gain. On reasoning-intensive retrieval benchmarks it beats strong baselines while keeping embedding inference cheap.
KV Cache Compression Through the Lens of Transform Coding
The key-value (KV) cache dominates memory in long-context inference, and existing quantization schemes minimize reconstruction error in the cache itself without asking how that error propagates through attention. Under a white-noise quantization model the authors prove the expected attention-aware distortion splits into additive key and value terms that factor across tokens and channels, which lets them apply transform coding and reverse water-filling from rate-distortion theory to allocate bits over a calibration set. The resulting Attention-Aware Transform Coding (AATC) reaches near-lossless accuracy at about 5.8× compression on Llama-3.1-8B-Instruct and Qwen-2.5-7B-Instruct across LongBench, RULER, GSM8K, MMLU-Pro, and MATH-500, while every baseline degrades in at least some settings.
FreeBalance: Pre-Routing Online Moe Load Balancing via Residual Workload Prediction
In distributed Mixture-of-Experts (MoE) inference, skewed routing leaves one rank overloaded and stalling everyone else, and existing online rebalancers can only act after each router has produced its decisions, putting expert weight migration squarely on the critical path. FreeBalance predicts the routing distribution in advance from cross-layer similarity of hidden representations in the residual stream, so expert migration can be planned early and overlapped with pre-routing computation such as attention, with a cost model capping the number of swaps so synchronization stays hidden inside the available window. Across models and datasets it cuts the max-to-mean rank load ratio by 32.8% and end-to-end prefill latency by 13.1%, hiding migration of an average 5.1 experts per layer that would otherwise cost about 8.5% of critical-path latency.
APTER: Adaptive Post-Training with Expert-Grounded Rubrics
Professional-domain deployments need models that respect domain constraints, cite critical evidence, and reason completely, but holistic preference tuning and outcome-level verification give no purchase on those requirements, and per-query generated rubrics tend to drift and skip essentials. APTER grounds rubrics in an expert criteria framework where each criterion names a stable professional capability, selecting and instantiating relevant criteria per query to produce executable supervision without reference answers. Rubric verdicts then serve double duty during reinforcement learning: as the optimization signal, and — when low scores are aggregated by criterion identifier — as a diagnosis of persistent weaknesses that triggers targeted supervised fine-tuning updates. Across three model generations, averages improve over the base models by up to 15.86 points on mathematical reasoning and 8.04 on medical question answering.
Multi-Objective Bayesian Optimization for Model Merging
Merging trained models in weight space avoids extra fine-tuning, but choosing merge coefficients is hard because evaluations are expensive, there are no gradients, and source capabilities trade off against each other. MOBO-Merge casts coefficient selection as black-box multi-objective optimization and uses Bayesian optimization to approximate the Pareto front within a fixed evaluation budget, independent of which merge operator is used. Testing Qwen3-4B and Llama-3.1-8B across two-model instruction-math and three-model instruction-math-code settings with Linear, SLERP, TIES, and block-wise operators, it beat random search on held-out hypervolume in 11 of 12 comparisons, with negligible benefit for one-dimensional linear interpolation and substantial gains for the higher-dimensional TIES, block-wise, and three-objective searches. No single operator dominated: TIES led three of four family-setting combinations while Block-Linear 4x won the Llama three-model merge.
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
On-policy distillation from a long-context reasoning teacher into a short-context student breaks down over tokenizer mismatch, distribution mismatch, runaway response length, and unstable training. SimpleOPD performs distillation in a shared text space, aligning only tokens that cover identical text spans under both tokenizers, and controls length by adding a student reference KL term plus masking the advantages of termination tokens such as </think> so the student grows its reasoning steadily instead of drifting and getting truncated. Transferring proof-reasoning ability from SU-01 to Qwen3, Qwen3.5, Intern-S2, GLM-4.7, and Gemma-4 students improved mathematical reasoning across both same-family and cross-family pairs, with Intern-S2-Preview gaining 21.2 points on ProofBench to reach 55.2 and surpassing Gemini-2.5-Pro, alongside gains on science benchmarks HLE and HiPhO that suggest transfer beyond the training domain.
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Mobius-v0 splits a transformer's usual entanglement of knowledge and computation into a globally shared Memory built from feed-forward layers that stores knowledge vectors, and multiple Reasoners built from self-attention that repeatedly query that memory using hidden states as both cache and carrier. Because knowledge lives in one shared store while reasoning is applied iteratively, the architecture compresses knowledge more tightly and reuses reasoning operators. A 7B model trained from scratch matched a 7B transformer baseline's downstream scores using 62.6% of the baseline's training data, and Intern-S2-Mobius, continually pretrained from Qwen3.5-35B, matched downstream quality while delivering nearly 4x end-to-end inference speedup.
AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs
Anchoring, the human bias in which an arbitrary reference number pulls a later numeric judgment toward itself, has been observed in language models, but prior studies test only a few ways of introducing the anchor and rarely separate obviously irrelevant anchors from plausible ones. AnchorBench evaluates multiple anchor pathways — including anchors injected through external context and retrieval-augmented generation — along an explicit relevance axis, across ten open-weight and four frontier API models. Susceptibility turns out to be strongly pathway-dependent, plausible anchors shift answers more than irrelevant ones when delivered through the stronger pathways, and influence fades as the anchor moves further from the evidence-supported answer. Notably, frontier models scoring above 95% on the anchor-free control condition remain vulnerable to plausible anchors, so task accuracy is no proxy for robustness.
Local and Global Regimes of Geometric Complexity in Language Model Representations
Intrinsic dimensionality is a common probe for how complex a language model's representations are, but it is unclear whether measured differences reflect language or artefacts of dataset construction. Holding everything else fixed and varying lexical diversity — the number of unique final tokens in a dataset — reveals a scale-dependent reversal: at low lexical diversity fewer unique final words give higher intrinsic dimension, while at high diversity the ordering flips. The authors derive an exact, parameter-free formula for the crossover point that matches the empirical transition at every scale tested, which both cautions against reading intrinsic dimension as a direct measure of complexity and describes an organising principle of the representation manifold.
DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding
Mixture-of-Experts (MoE) models serving latency-sensitive workloads such as coding assistants run at small batch sizes, where inference becomes memory-bound and the dominant cost is streaming expert weights rather than computing with them. DeaMoE restructures the expert layer into "departments" whose member experts share most of their parameters and keep only a small private residual, paired with a two-stage router that avoids loading redundant weights. Per-step loaded weights fall by up to 50.9%, giving up to 1.33x end-to-end time-per-output-token (TPOT) speedup for a pretrained 7B model on an A40, and microbenchmarks on DeepSeek-V3 show peak speedups of 2.00x on A40 and 1.97x on H100.
More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It
Power Sampling sharpens a language model's distribution over whole generation trajectories, a verifier-free way to push probability mass toward correct reasoning paths that is often proposed as a front end for other inference-time methods. Testing it with self-consistency reveals a paradox: it can increase the mass on correct trajectories while cutting downstream accuracy by as much as 18.5 percentage points across models and reasoning benchmarks. The authors attribute this to a dose mismatch, where a single fixed exponent changes different problems by wildly different amounts, and a coverage mismatch, where global sharpening collapses onto a few dominant paths — so a high pass@k can coexist with the loss of the broad path support that aggregation and search need. Replacing uniform trajectory exponentiation with a deformation-controlled, support-preserving target that calibrates sharpening per problem reverses the losses and, at equal budget with weighted self-consistency, beats standard multi-sample inference.
Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations
Benchmark runs typically sample every item the same fixed number of times, spending compute on items whose scores stopped moving many epochs ago. optstop reframes evaluation as sequential measurement built on hierarchical Bayesian inference: sampling continues where uncertainty is still high and halts where estimates are precise or stable, with support for binary, ordinal, and continuous outcomes, no need for a calibrated item bank, live or retrospective operation, and a safeguard that samples more cautiously near zero performance where rare successes carry the most information. On an illustrative 200-item, 10-epoch evaluation it eliminates 57% to 97% of planned trials across nine validation settings while reaching the same overall conclusions, with savings depending on how the evaluation is designed.
Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation
Summarization metrics measure overall quality or specific properties like factuality, but not whether a summary actually serves a particular reader — a biomedical researcher and a family physician want different things from the same vaccine paper, and a short query rarely captures that gap. The argument is that a reader's persona, meaning their role and expertise, is more stable than any single query and recovers the missing context, so metrics should be tested for sensitivity to both informational and persona differences. Perturbation tests show that popular metrics including strong LLM-as-judge scorers fail basic checks on informational content, and an expert human study of persona-conditioned preferences finds that traditional and LLM-based metrics alike agree poorly with human judgments of information satisfaction.
You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model
A frozen language model both under-uses evidence already present in its residual stream and fails to notice when the input cannot support an answer, so it confabulates; two known fixes — a conditional steering probe that writes to mid-stack layers, and a zero-shot sufficiency direction that reads the stream to trigger abstention — interfere when combined, since steering shifts the very state the direction inspects. YOPO keeps the direction fixed and trains a small network to reconstruct the pre-steering residual from the steered one using mean-squared error on paired activations and no sufficiency labels, then reads the direction on that reconstruction, so answering, steering, and abstaining all happen in a single forward pass. On frozen Qwen2.5 backbones at 1.5B, 3B, and 7B, three-way accuracy more than doubles the frozen baseline (0.375 to 0.798 on alphaNLI at 1.5B) and the single pass beats the two-pass reference at every scale and across ten backbones from six families. A source-side audit caught a surface artifact leaking in the authors' own alphaNLI construction, so the architectural claims rest on native-label replications with SQuAD2, RepLiQA, and MuSiQue.
Approximate Muon with low-rank adapters
The Muon optimizer helps during pretraining but is seldom used for parameter-efficient fine-tuning, largely because LoRA's low-rank parameterization makes it mathematically impossible to orthogonalize the resulting weight update. The fix approximates a relaxed Muon objective in the low-rank setting via linearization followed by least squares, with an implementation that needs only matrix multiplications instead of costlier decomposition routines. Across supervised fine-tuning and a ReLoRA pretraining run, sMuon performs favorably, though the authors are explicit that results vary by model and evaluation and that the overall gain from Muon-style low-rank fine-tuning is moderate rather than dramatic.
Split the Labor: Separating Evidence Interpretation from Decision Aggregation
Systems that ask a language model to draw a conclusion from many sources typically concatenate everything into one prompt, conflating two jobs with different needs: interpreting a single source rewards model capacity and context, while combining interpretations rewards fixed arithmetic, cross-instance comparability, and the option to return nothing. Splitting them turns the design problem into the interface, here a four-field evidence tuple of hypothesis, reliability bucket, rationale, and provenance — and exposes a failure mode the authors name count-scale drift, where thresholding a sum of unnormalized weights is posterior thresholding at an operating point that slides with how many sources were consulted, so no single threshold reconciles vote order with posterior order when source reliabilities differ. Pooling calibrated log-likelihood ratios fixes both problems arithmetically rather than architecturally, and applies equally to score-summing triage engines, diagnostic panels, and additive detectors. Instantiated twice on one longitudinal corpus, the separation helps in both settings, with a small sequence encoder on an auxiliary objective plus a tree ensemble carrying a censored survival loss reaching 0.921 AUPRC against 0.805 for a hand-crafted baseline; the paper also states five falsifying predictions, three negative results, and which comparisons remain confounded.
Handover of In-Context Learning State Across Session Boundaries
When a long-running task outgrows a model's context window, the application restarts, or another agent takes over, something has to decide what information carries into the next session. The work formalizes this as transferring a task-relative in-context learning (ICL) state, separating exact recovery of the earlier text from preservation of the target predictive distribution, and shows that under an exogeneity condition predictive equivalence characterizes the coarsest sufficient handover and yields a fixed-length bit requirement. It proposes a three-part record — decisions and constraints stored verbatim, task-justified summary statistics for repeated evidence, and raw observations whose effect those statistics fail to capture — and quantifies the penalty for writing the record before the downstream query is known. Gaussian linear regression admits an exact finite-dimensional handover with finite-bit perturbation bounds, while nonparametric regression yields matching upper and lower bounds linking memory size to squared prediction error.
2 more specialized papers
Agents 39
Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation
Scaling agent evaluation usually means swapping an executable environment reward for a second language model as judge, and such judges — whether hand-written like G-Eval or fine-tuned — tend to credit fluent but unsuccessful trajectories as successes. RubricForge instead evolves the text of a judging rubric against a small set of ground-truth-labeled trajectories to maximize agreement with the real environment reward, then freezes it and applies it in one model call with no environment access, leaving a human-readable artifact whose verdicts trace to named criteria. Using a single 7B model as both agent and judge on tau-bench and WebShop, aggregate agreement with G-Eval is statistically indistinguishable, but the false-pass rate on failed trajectories drops roughly by half, 0.115 versus 0.173 on tau-bench — the quantity the authors argue matters, since a false pass ships a broken agent while a false fail only costs a retry.
Does a Language Server Save Tokens for Coding Agents? A Measurement Methodology and Preliminary Study
Coding agents spend most of their context budget on retrieval, and the claim that typed semantic retrieval through the Language Server Protocol (LSP) is more token-efficient than lexical grep is, the authors note, asserted almost everywhere and measured almost nowhere. They formalize the question with a tokens-to-success metric, specify a five-arm ablation isolating semantic retrieval from confounds, and run a preliminary study on Python and TypeScript repositories with Claude Opus 4.8, Sonnet 4.6 and Haiku 4.5. The answer is conditional and usually negative: on symbol-named localization the LSP costs 6% to 118% more tokens and agents ignore it even when it is free, saving tokens only for the weakest model on reference-completeness tasks. On multi-file renames scored by real test execution, grep succeeds perfectly while a location-only LSP fails three-quarters of them, since a rename must touch comments and strings that semantic references exclude — pointing to an adaptive router keyed on task class, model capability and lexical noise rather than LSP-always.
Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems
When an agentic system retries after a failed answer, the extra tokens open a gap between a model's advertised per-token price and what the workflow really costs — token inflation, defined as the ratio of true workflow cost to single-call cost and measured as high as 4.25× for a 7B model on multi-hop question answering. InflationAgent routes on that predicted true cost rather than list price, using CoT Branching Entropy, a pre-execution difficulty signal computed from local inference alone (AUROC 0.887), and selecting models by expected accuracy divided by predicted cost, with a policy that discards failed chains before escalating. On GSM8K under a fixed budget it reaches 94.7% accuracy versus 91.0% for FrugalGPT while using 31% fewer tokens, and forwarding a failed reasoning chain to GPT-4o is shown to cost up to 34.8 percentage points of accuracy, supporting the fresh-escalation design.
Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents
Tool-using agents that modify local state, keep persistent memory and speak external protocols bring risks of over-privileged actions, weak auditability, prompt injection, tool poisoning and uncontrolled side effects. Agentao is a local-first runtime that separates model-generated action proposals from host-authorized execution, layering host-facing surfaces, a host contract, a runtime core, a permission-mediated tool system, and subsystems for memory, replay, plugins, skills, sub-agents and protocol integration. The authors are explicit that no formal safety guarantees follow; the contribution is making permissions, state, protocol boundaries and execution traces explicit runtime abstractions, with the threat model, governance model and structured event interface described and the code released publicly.
Measuring Cross-Task Behavioral Consistency in Language Model Agents
Agent evaluations report success rates, which say whether a system solved a task but nothing about whether it approaches different tasks the same way. The Behavioral Consistency Metric (BCM) trains a model to predict task success from behavioral features of execution traces, extracts a per-trajectory feature-attribution vector, and averages pairwise similarity of those vectors within an agent system. Over roughly 9,000 trajectories from six language model agents on software engineering tasks, within-task reproducibility and cross-task consistency turned out to be separate axes: some systems repeat themselves faithfully on retries of one task yet have no stable strategy across tasks, a split invisible to prior same-task reproducibility measures, and consistency did not track success rate.
MobileMem: Learning from a Year of Mobile Experiences
Personal assistants that accumulate knowledge about a user over months need benchmarks built from heterogeneous, multimodal, evolving personal data, which existing long-term memory evaluations do not provide. MobileMem supplies both a benchmark and framework for on-device memory, using a knowledge-grounded synthesis pipeline to turn a year-scale collection of real user-app sessions into coherent, temporally consistent long-horizon trajectories. It offers matched text-only and multimodal settings that probe multi-hop and temporal reasoning, updating superseded knowledge, and inferring preferences the user never stated, framing memory as accumulated experience rather than a retrieval index over isolated facts.
Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
Systems that pair a language model with retrieval or memory so it improves from feedback without retraining are hard to evaluate in security operations, where benchmark labels are scarce, stale, and unrepresentative. The proposed alternative dispenses with labels entirely: a stronger teacher model supplies sparsely sampled corrections to a smaller student running the harness, and the harness is scored by how far the student converges toward the teacher over time. Across security tasks, model families, and harness designs, teacher-relative lift correlated with improvement measured against a held-out gold standard, while LLM-as-a-judge comparisons between similarly capable models produced no usable signal at all.
SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data
Natural-language front ends to enterprise databases have to turn vague requests into queries that are correct, policy-compliant, cheap, and repeatable, and SemPlan compares four architectures for doing so on a deterministic synthetic benchmark of 1,800 English and Brazilian Portuguese cases (1,200 frozen for scientific evaluation). The contenders are direct SQL generation, a bounded tool-agent, structured semantic-request generation followed by deterministic planning and execution, and a clarification-capable stateful variant of the latter. Across 4,800 records answer correctness stayed low for every architecture — 22.25%, 22.58%, 25.67%, and 24.25% respectively — with the deterministic planner most correct, direct SQL safest on policy compliance, and the clarification variant cheapest, supporting a trade-off reading in which added structure shifts failure modes rather than eliminating them.
Ontology-Grounded Project Memory for Coding Agents
As coding agents generate more of a project's code, the reasoning behind those changes becomes hard to track. MOOSEDev stores architectural decisions, lessons, constraints, and rationales as typed records in a knowledge graph with lifecycle status, provenance, and supersession links, exposed to agents through a Model Context Protocol (MCP) interface and queried by a neurosymbolic engine that treats the symbolic layer as the primary reasoning substrate. On a neutral public corpus of 835 records, it returned essentially the full expected answer set (0.98-1.00) on supersession, set-completeness, and negation questions, versus 6% to 27% for a production vector-memory baseline, while relevance recall and token cost were comparable between the two.
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
Agents in the ReAct paradigm think only during the Thought phase and sit idle while an action is serialized and the environment responds, leaving a recurring reasoning idle window unused. Second Thought is a training-free inference framework that forks four auxiliary reasoning branches the moment a Thought phase ends, decodes them concurrently with the main loop, and merges their output back when the observation arrives, moving the extra reasoning off the sequential critical path. Across three agentic benchmarks and three reasoning models it lowered average turn count in all nine model-benchmark pairs and cut main-thread decoding by up to 43% with Pass@1 unchanged in seven pairs; against a compute-matched control that spends the same budget on the main thread, it achieved strictly higher Pass@1 with 1.3 to 3.2 times less sequential decoding.
CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA
Multi-agent retrieval pipelines usually debate a finished draft rather than the individual assertions inside it, and treat every retrieved modality as equally trustworthy, so hallucinations introduced between agents go unnoticed until the final text. CLAIR-Fin splits each question into atomic claims tracked in a typed ledger, conditions how much a claim trusts each evidence modality on the claim's type, checks grounding at the hand-off from drafting to adversarial review, and routes contested claims into a debate whose depth scales with what the debate turns up. On BB-FinQA-X, a 500-question cross-modal set built from Bangladesh Bank annual reports, faithfulness rises from 0.780 to 0.889 against a single-pass retrieval-augmented generation baseline while the system abstains on 5.4% of questions with insufficient evidence, also beating HyDE and Graph-RAG (both at most 0.874).
Simulation-Aware In-Context Policy Improvement for LLM-Aided Analog Layout Refinement
Analog integrated-circuit layout still depends on experts hand-tuning optimization parameters between rounds of parasitic extraction and post-layout simulation, and Bayesian optimization needs hundreds to thousands of such evaluations to help, which is far too many at layout level. The proposed multi-agent framework performs in-context policy improvement: agents run an act-observe-reflect loop over compact structured representations of the layout, updating the parameters exposed by an analog layout generator between simulations. On real analog circuits it improves post-layout performance over both the generator's built-in heuristics and Bayesian-optimization tuning using only tens of post-layout simulations.
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL
Agents acting for a user often face a counterpart with opposing goals, and the agreeable dispositions of a helpful assistant make for a bad delegate: frontier models volunteer private information and concede early. SocialRL trains social reasoning directly in a 4B-parameter model across six negotiation and scheduling domains including Deal-or-No-Deal, CaSiNo, Craigslist, and a job-interview setting, with every policy evaluated on all six. In-domain training closes 73–122% of the gap from baseline to frontier, and consolidating the specialists through cascade reinforcement learning and multi-teacher on-policy distillation yields a single 4B model averaging 0.627 utility across all six environments, matching or beating GPT-4.1, GPT-5.1, and GPT-5.2. Transfer follows game structure — structurally paired games lift each other while isolated ones transfer nothing — and distilling theory-of-mind traces rather than actions alone helps everywhere, though only next-action prediction correlates with negotiation outcomes.
AdsWorldEngine: A Self-Evolving Conversational Advertising Agent through Orchestrator and Tool Coevolution
Ads inside multi-turn assistant conversations must infer commercial intent from the query, the assistant's reply, and dialogue history, and also judge whether an ad would help or annoy. AdsWorldEngine splits this across an Opportunity Gate that decides whether to show anything, an Orchestrator that generates intents, calls advertising tools, and assembles a top-three slate, and an Evaluator that scores delivered ads offline. The Orchestrator is trained by supervised fine-tuning plus agentic reinforcement learning, and its high- and low-reward rollouts then become preference data for training the tools themselves, creating a loop that improves tool use and the tools together; subjective calls are handled by judgment models distilled from human labels with reflection-filtered rationales and a cost-sensitive GRPO variant. Offline diversity rises 60% and relevance 80% over the production system, and an online A/B test shows 22% higher revenue per mille and 74% more ad coverage.
Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model
Coding agents are benchmarked as models but shipped as systems whose reliability depends on the harness, execution environment, retrieval, memory and state handling, permissions, review interfaces, and resource allocation. This monograph synthesizes 164 scholarly works, 100 practitioner records, 29 benchmark records, and 17 operated-system case records into a framework that treats evaluation and operation as a dependency chain, arguing that many apparent model failures actually originate elsewhere in the system, and that gains at one layer often fail to show up in end-to-end outcomes. It contributes a versioned catalog of 206 reliability records, an evidence ledger, runnable evaluation and reliability protocols, and five reusable agent skills, while noting that the review is structured rather than exhaustive and that evidence strength varies by topic.
MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends
Most agent-memory benchmarks probe after-the-fact recall rather than whether stored memory actually carries an agent through interdependent multi-session tasks, which is what MemoryArena measures. Holding the agent framework, model alias, task samples, and scoring code fixed, the study swaps only the memory backend, comparing MemoryLake against Mem0, vector retrieval over text-embedding-3-small, and a long-context control across five domains. MemoryLake leads on mathematics, physics, and progressive retrieval and averages 20.5% success versus 13.6% for the best comparator, though every system scores zero on travel planning and near-zero on web shopping, and the authors stress that sample sizes are small, confidence intervals overlap, and no significance tests were run.
Agentic Transaction: Towards ACID-Compliant Agent Systems
As language model agents move from chat into long-horizon work over persistent workspaces, they hit the same problems transactional databases were built to solve: reliable execution, consistent outcomes, safe concurrency, and durable state. The proposed framework reinterprets atomicity, consistency, isolation, and durability as semantic guarantees for agent execution, and instantiates them in a data agent using transactional exploration-execution-validation cycles, transactional skill hubs, confidence-divergence validation, semantic dependency-aware isolation, and transaction-aware state management. On widely used benchmarks the system reports a 10.6% improvement over state-of-the-art agents including Claude Code.
When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict
Agents that carry personal memory across sessions inevitably accumulate contradictions, and when a query provides no context, timing, or source authority to resolve them, picking one memory as definitive turns genuine uncertainty into a confidently wrong action. TANGLE is a benchmark of 541 deliberately unresolvable conflicts across 40 personas, split into context-partitioned, behavior-oscillation, and source-contradiction types, scored on conflict perception, causal reasoning, confidence calibration, clarification seeking, and memory faithfulness under both a curated-memory oracle track and an end-to-end extraction pipeline. Models notice conflicts far more reliably than they calibrate their actions or ask targeted clarifying questions, and pipeline extraction discards the conflict-bearing relations downstream reasoning needs, motivating a policy that adapts the response to the specific conflict rather than applying fixed rules.
AI Research Preference Models
Research agents can write candidate machine learning experiments in minutes but need hours to days of GPU time to evaluate each one, so progress depends on how a fixed execution budget is allocated across proposals. The approach builds preference models (RPMs) from frozen pretrained language models with no task-specific training, in two forms: an inference-only variant that reasons over candidate plans, code, and previously executed solutions, and an agentic variant that first runs small-scale pilot experiments. Wired into the AIRA-dojo search agent and measured on AIRS-Bench, the two variants lift the average normalized score from 0.684 to 0.711 and 0.729, and reach the unguided agent's 24-hour performance in roughly 15 hours using under two-thirds of its execution budget, with new state-of-the-art results on two benchmark tasks.
HELIX: Model-Harness Co-evolution for Recursive Self-Improvement
An interactive agent acts through a runtime harness that governs context, tools, control flow, and stopping, which means the harness shapes both what a model can do and the trajectories it later learns from. HELIX is a source-traceable substrate that decomposes agent systems into typed ports, reusable atoms, recipes, product shells, and runtime policies, making each intervention explicit and auditable while retaining trajectories, test outcomes, and provenance, so harnesses can be evolved for a fixed model and then rebuilt as the model improves. In one evolution round on code repair evaluated with the SWE-bench evaluator, a 65-candidate portfolio found a fixed harness improving task coverage by 4.0% over Pi, while the full portfolio exposed up to 58.0% more verified coverage through complementary sibling behavior, and a 200-slot sibling slice yielded 438 verified supervised, critic, filter, and preference training records.
MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning
Answering what happens before, after, or across stages of a tens-of-minutes surgical video requires grounding a question in evidence spread across time, which one-shot vision-language models lose by compressing the procedure into a context window, while trained video agents are data-hungry and transfer poorly to unseen surgery. MedClaw separates reasoning from perception: a text-only orchestrator plans what evidence to gather and issues an auditable sequence of tool calls that frozen vision-language sub-agents execute over pixels by viewing, cropping, inspecting frames, and retrieving external knowledge, and a gradient-free, reward-gated Heuristic Skill Distillation loop mines low-scoring traces and keeps a candidate skill only if it raises validation reward, yielding reusable retrieval skills such as directed re-look. Because it grows an external skill library instead of tuning weights, the loop adapts from roughly 100 labeled examples, and on the new MedClawBench of 1,123 doctor-grounded questions over long neurosurgery recordings plus a held-out lecture-video split it beats one-shot models and general video-agent frameworks across all four evaluation dimensions.
Agent-Orchestration in Autonomous Chip Design
Framed as a position piece on where tool-using language-model agents fit in integrated circuit design, the work proposes treating a chip-design system not as a single model but as a large AI-organization of coordinated agents. The argument centers on what kind of artificial intelligence the industry actually needs given how specialized and interdependent the design flow is. The contribution is the organizational framing itself rather than a benchmarked system.
Demystifying Agent Skills: Why They Work-Until They Don't
Agent skills — structured packages of procedural knowledge injected at inference time — are widely reported to raise task success, but little is known about the mechanism or the failure modes. Controlled experiments across several benchmarks, agent harnesses, and language models isolate the effects of skill representation, outcome annotation, retrieval difficulty, and cross-framework robustness, and a paired-trajectory contrastive study normalizes 8,135 trial records into a taxonomy of twelve skill-use modes. Procedural anchoring — stabilizing the sequence of actions rather than supplying missing facts — accounts for 65.7% of cases where skills help, versus 4.5% for explicit knowledge injection, and skills beat workflow memory by 6.06 points in matched comparisons. Retrieval turns out to be a separate bottleneck: as the skill pool grows from 5 to 100, the precision of skills actually used falls from 29.6% to 3.3%, though exact ground-truth skill selection proves neither necessary nor sufficient for downstream success.
MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation
Merchant-deployed conversational shopping assistants must recommend only from a fixed catalog, honor hard constraints like budget and brand exclusions, and keep those constraints straight across turns — reliability requirements that prompt-only language-model systems handle poorly. MACS splits the work: the language model interprets requests, elicits preferences, and writes responses, while product retrieval, hard-constraint filtering, brand exclusion, and progressive constraint relaxation run deterministically in a merchant agent backed by a session-persistent preference layer. On a 140-query single-turn benchmark it reaches an 87.1% pass rate with perfect brand compliance, and on a 10-scenario multi-turn benchmark 72% macro Pass@5 with zero constraint drift versus 56% and 52% for catalog-bound GPT and Gemini prompting, with the largest margins on exclusion reversal and accumulated constraints.
A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents
Long-running LLM agents can silently drift away from their assigned task and cause irreversible side effects on external systems, and prompt-level guardrails offer no step-level detection, risk scoring, or recovery decision. Because the main executor is usually a large model that cannot be retrained per deployment, this work trains a separate small language model with reinforcement learning to occupy each node of an external recovery graph — drift classification, operation detection, risk evaluation, final decision — emitting XML-structured reasoning tailored to that role, with rewards combining schema and length rules with an LLM-as-judge score for semantic quality. On the public AppWorld benchmark the small model generally makes correct recovery decisions when given information about the suspected drift onset, and reliably respects the prescribed output schema at every node.
Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions
Mobile GUI agents built on multimodal large language models (MLLMs) mostly react to explicit instructions, with no pipeline for inferring what a user is trying to do, predicting the next intention, and acting on it. Act2Intention supplies both a benchmark — 72,511 intentions and over 700,000 actions across 52 apps, collected and validated as continuous intention-action trajectories — and an agent that chains proactive intention understanding, personalized prediction, and experience-guided execution. Supervised fine-tuning on the benchmark yields absolute gains of +32.0, +10.25, and +6.9 points on understanding accuracy, prediction accuracy, and step success rate over the same framework without fine-tuning.
AutoSchema: Live Schema Grounding for Agentic Text-to-Sparql over Heterogeneous Knowledge Graphs
Life science knowledge graphs each expose their own SPARQL schema and identifiers, so agents querying them typically depend on hand-curated metadata files that require language-model drafting plus manual review to maintain. AutoSchema instead does live schema grounding with no training: it inspects endpoints directly, maps entity names in a question to graph identifiers, explores relation paths, and discovers cross-resource links during iterative query construction. Against TogoMCP as the curated-file baseline, it raises mean factoid accuracy on biomedical knowledge-graph question answering while using fewer tool calls and exhausting its iteration budget less often, with consistent gains on a longitudinal BioASQ Task B evaluation and preliminary evidence of transfer to an undocumented chemistry RDF graph.
Polaris : Multi Agentic System for Conversational Enterprise Analytics
Enterprise analytics stalls when business questions require chaining query generation, visualization, and explanation across systems that non-specialists cannot address directly. Polaris is a supervisor-led multi-agent framework whose Dynamic Task Coordination layer treats agent-task assignment as adaptive bipartite matching, letting the supervisor reassign and recover across specialized querying, visualization, and reasoning agents at runtime; these agents follow a reason-first ReAct loop so answers include the underlying explanation rather than just retrieved rows. Evaluation on structured enterprise datasets reports high semantic fidelity and answer relevancy, though the paper describes results qualitatively rather than against a public benchmark.
TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments
Time series question-answering benchmarks are built from frozen snapshots, so they never test whether an agent respects data cutoffs or revises conclusions when a new release changes the evidence. TimeSage-EV tracks 60 real institutional scenarios across 6 domains with 1,485 scenario-period question-answer pairs spanning February 2023 to May 2026 and monthly, weekly, daily, and irregular release cadences; at each period an agent sees the series and source reports while the withheld next release supplies ground truth for state identification, summarization, and outlook reasoning. Experiments with frontier LLM agents and TimeSage-1.0, a self-evolving agent that accumulates a reusable analytical skill library, show large gaps between model tiers and recurring failures in temporal validity, use of exogenous context, and adaptation to new releases. The benchmark is released with monthly updates, code, a leaderboard, and failure-mode analyses.
Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents
Agents built on language models tend to act on what they already know rather than deliberately gathering information that would improve later decisions, a capability the authors call proactive exploration. SAFARI attacks two identified bottlenecks: standard expert demonstrations suffer hindsight bias because the demonstrator already knew the answer, so an Exploratory Data Construction stage synthesizes trajectories rich in genuine information-seeking; and reward signals conflate useful exploration with aimless wandering, so reinforcement learning is guided by contrastive trajectory pairs that separate productive exploration from redundant wandering. Experiments across environments report gains from both components along with analysis of what proactive exploration looks like in practice, and the code is released publicly.
ATLAS: Discovering Agent Strategies through LLM-Guided Abstraction and Automata Learning
Agents driven by language models are now run on software testing and cybersecurity tasks, but evaluation usually stops at task success and raw execution traces, which say little about the strategy an agent actually followed. ATLAS (Automata Learning for Agent Trajectory Analysis and Strategy Discovery) abstracts trajectories into symbolic events and applies automata learning to infer finite-state models of agent-environment interaction that expose recurring behaviours, decision points, successful completion paths, and failure loops. Applied to a penetration-testing agent across 12 vulnerable machines, the learned automata surface high-level exploitation strategies that are hard to read off raw traces, and the authors further demonstrate transferring the extracted symbolic model from a frontier model to a compact one, plus model transformations that yield concise behavioural explanations.
ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond
Autonomous research agents stall over long horizons because they lack ways to preserve continuity, back out of dead ends, and spend compute where it is paying off. ScienceFlow structures a research run as segments grounded in executable workspaces, representing progress as recoverable executable states; transitions are governed by Executable-State Transition through Re-Anchoring (ESTRA), which picks either the live state or an archived one as the next anchor and decides whether to continue or redirect, while an evidence-aware controller allocates compute to jobs based on availability, remaining budget, and validated progress. Across machine learning, scientific modelling, and mathematical optimization tasks, it reaches 70.22% Any-Medal on the full MLE-bench within a 24-hour budget, 4.92 percentage points above the prior best reported result.
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages
Multi-agent reasoning pipelines usually decide which intermediate messages to keep by proxies for correctness such as agreement, confidence, or automated scoring, which assumes a wrong message is not worth passing on. Diverse Hypothesis Deliberation tests that assumption by caching five independently generated messages and replaying the same downstream solver twice, once with a given message visible and once with it hidden, so the causal effect of each message on final correctness can be measured directly. Across five mathematics and science benchmarks with gpt-oss-120b and gemma-4-31B-it, messages carrying wrong answers but useful decompositions or constraints appear in every benchmark-model pairing, and more than four in ten of the correctness-flipping wrong-answer messages flip it in the helpful direction. Passing the complete message beat passing only its reasoning, which in turn beat passing only its answer, and the resulting replay labels can be reused to train when agents should listen to each other.
AgentRewind: Recoverable Execution for Long-Horizon LLM Agents
When a language-model agent works over a long horizon, an early mistake contaminates both its context and the environment state, and later actions often cannot undo it; existing defenses concentrate on preventing errors through planning and safety checks rather than recovering from them. AgentRewind is a runtime framework that records aligned checkpoints of the agent's context and a controlled environment, letting the agent roll back to an earlier state and retry while carrying forward what it learned from the failed attempt. The authors also build MettleBench, a benchmark of long-horizon engineering assignments made of chained requirements that scores both completion and partial checklist progress. Across several models, execution strategies, and agent harnesses, rollback-and-resume raises both task success rate and average checklist progress over the baselines.
The Past and Future of AI Scientists
Machines that originate hypotheses, deduce consequences, design and run experiments, and revise beliefs have existed since Adam made the first novel machine discovery through physical experimentation and Eve established the self-driving-laboratory architecture; foundation models, autonomous agents, and lab robotics now make far more general systems buildable. Surveying that history and what comes next, the authors argue the open problem is no longer automating individual components of science, which is already possible, but integrating them — combining neural learning with logic, probability, mathematics, causal reasoning, simulation, experimental design, robotics, and formal scientific records. They assess progress against the Nobel Turing Challenge's 2050 target for automating Nobel-quality discovery and judge it ahead of schedule.
PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments
Self-evolving agents are usually measured under fixed execution conditions, so nothing tests whether they can recover once the environment changes underneath them. PACE-Bench supplies 144 source-to-target adaptation pairs across six simulated physics domains, where a code-driven design that works in the source environment breaks in a mutated target with the same goal and interface, and the agent must repair it using sandbox diagnostics within a limited attempt budget. Across ten methods from four paradigms the benchmark stays far from saturated — Reflexion with Qwen3-14B solves only 35.9% of pairs, and GPT-5.5 reaches 66.7% on the Statics subset alone. Simulator-grounded reflection beats unverified self-revision, memory tends to anchor agents to their initial design, and even disclosing the exact physical change does not lift the ceiling, suggesting the bottleneck is redesigning mechanisms rather than inferring parameters.
Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports
Generative models used to synthesize technical write-ups tend to produce ungrounded claims, and they rarely handle non-text content well. Wyvern is a multi-agent pipeline that assembles technical reports combining text, figures, and tables with supporting citations, adding an automatic claim-revision stage aimed at keeping assertions tied to sources. In a human study, evaluators judged its figures more informative than a recent baseline in 87% of comparisons and preferred its reports over three alternatives in 63% to 100% of cases, while automatic scoring showed up to 2.3× higher citation recall and 1.6× higher citation precision than the baselines.
SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning
Real workbooks carry implicit cross-table associations, fine-grained column dependencies, and spatial layout that large language models lose when a sheet is flattened into a linear string. SheetCompass instead builds explicit hierarchical relation graphs capturing structure within and across worksheets, and pairs them with a memory component that retains task-relevant state as an agent works. The claim is that preserving intra-sheet boundaries and inter-sheet semantics lets agents use the global spatial context human analysts rely on when reasoning over and automating complex spreadsheets.
Twin: Playing an Unknown Game with a Test-Time Digital Twin
Playing a game whose rules and goal are hidden usually means hand-engineering a world model per task; Twin instead has a frontier coding agent write an executable world model from interaction alone, on continual-learning tasks such as ARC-AGI-3. A harness enforces replay validation — no action is taken until the program reproduces every previously observed transition — and each prediction mismatch becomes a counterexample that drives a repair of the program. The system clears 179 of 183 levels (97.8%), beats human efficiency on 158 of them, and infers the goal before receiving any reward on 156; on the benchmark's human-referenced 0–100 score the same base model goes from 7.8% played directly and 61.1% with an off-the-shelf harness to 93.3% with the twin world model. The authors conclude that constructing a usable world model was easier than expected, while identifying the right goal is the harder half.
Other 29
The Architect: Interactive Visualization of Deep Learning Mathematics Directly in Microsoft Excel
Deep learning libraries hide the arithmetic behind function calls, and most visualization tools stop at architecture diagrams or training summaries. The Architect takes a compact table describing a neural network and generates a Microsoft Excel workbook in which the full forward pass — and, on request, backpropagation and parameter updates — appears as live spreadsheet formulas that recompute automatically when the user edits inputs, weights, labels or hyperparameters. Matrices, activations, losses and gradients become inspectable spreadsheet regions, aligned PyTorch snippets connect formulas to code, and the report walks through arithmetic tracing, learning-rate exploration, and diagnosis of dying ReLU units and vanishing gradients.
AI Evaluation Should Work With Humans
A position argument that the prevailing evaluation paradigm — measuring superhuman autonomous performance — implicitly aims AI development at replacing people rather than complementing them, and is steering the field in the wrong direction. The proposed pivot is to evaluate human-AI teams rather than models in isolation, on the grounds that team-level metrics would push development toward systems that genuinely complement human capabilities and produce better societal outcomes.
Hard Cases, Bad Labels: Testing Error Exposure and Error Location in Uncertainty Sampling Under Bounded Label Noise
Active learning picks the examples a model is least sure about, but those same examples tend to be the ones human annotators get wrong, raising the question of whether uncertainty sampling suffers because it collects more bad labels or because bad labels in hard regions hurt disproportionately. Margin-based selection is compared against random sampling on three public binary tabular datasets under clean labels, random classification noise, and bounded difficulty-dependent noise, using 100 paired seeds, nine noise rates, and budgets from 20 to 120 labels with regularization re-tuned at each budget. Uncertainty sampling gained 1.09 to 1.77 percentage points of normalized balanced-accuracy area under the learning curve with clean labels, and exposure-matched controls found no evidence for a universal extra penalty from errors being concentrated in difficult regions — the apparent robustness varied by dataset, budget, noise structure, and which metric was used.
Asymmetric Discourse Homogenization and Shared Language Technology: Evidence from Reddit
Using six million Reddit comments from two cross-partisan forums between 2019 and 2025, the analysis finds an ideologically asymmetric break around late 2022: conservative users' prior trend toward more diverse political discourse was interrupted while progressive users showed no comparable change, a pattern that holds across interrupted time series, difference-in-differences, regression discontinuity in time, and propensity-score matching. A permutation test over 2,377 candidate cutoff dates places the ChatGPT release at an unremarkable 49.8th percentile, indicating gradual buildup rather than a single break, though a continuous cumulative index of exposure across seven model releases stays significant under a quadratic trend. Restricting to authors active throughout the window makes the effect vanish, so the homogenization looks community-level and ecological rather than driven by individual users adopting AI writing tools, with concurrent secular change not fully ruled out.
Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection
Physiology constrains how human speech unfolds over time, suggesting synthetic speech should violate those temporal regularities in detectable ways. A causal LSTM next-frame predictor trained only on bonafide speech over features from the Wav2Vec2-Large-AntiDeepfake backbone tests this directly, with a global-average-pooling baseline on identical features isolating the contribution of temporal modeling, plus an optional supervised stage that fits a multilayer perceptron on the frozen LSTM states. Across ASVspoof 2019/2021, Codecfake, In-the-Wild, MLAAD-EN, and Deepfake-Eval-2024, the system reaches a best published 0.75% equal error rate on ASVspoof 2021, and the bonafide-only stage alone beats the published supervised baseline from the same backbone on Deepfake-Eval-2024. Static and dynamic features tie on near-domain data, while trajectory dynamics provide the gains on harder cross-corpus benchmarks.
Emergent Models: Intelligence from Tiny Substrates
Emergent Models treat modeling not as learning a closed-form input-output map but as coaxing computational behavior out of simple open-ended substrates such as cellular automata, which iterate a fixed local rule over a latent space for an adaptive number of steps with an interface connecting latent state to external signals, trained by evolutionary search. The hypothesis is that some such systems are biased toward global generalization, capturing the rule that generated the data over its full domain and extrapolating past the training range. Theoretically, certain instances are proven latent-universal: with the update rule and interface fixed, varying only the initial latent state realizes any partial computable function; empirically, a zoo of minimal discrete and continuous substrates with tens to hundreds of parameters extrapolates exactly on simple arithmetic and supports control and online adaptation, while exposing clear limitations. The stated aim is to widen the design space beyond differentiable feed-forward maps rather than to offer a competitive architecture.
Attributing Preprocessing Invariance in Spectral Foundation Models
Spectral foundation models are often credited with learning invariance to preprocessing because a classifier trained under one pipeline still works under another, but these models normalize inputs before any learned weight is applied. Analyzing a Raman spectroscopy foundation model, the authors show that per-spectrum normalization collapses two differently preprocessed spectra to the same vector exactly when one is a positive multiple of the other plus a constant — a form many standard preprocessing steps take — so the encoder never sees a difference to be invariant to. Measured against its own parameter-free normalization on six Raman datasets, the model shows no measurable gain, and a controlled experiment confirms it only learns to ignore a transformation that actually reaches it. A numerical test identifies which transformations a given normalization already removes; surveying released systems across five modalities, most normalizations remove such transformations, and replication on two systems claiming learned invariance again found no gain.
22 more specialized papers
- Robust XGBoosting for Regression Iris Arag\'on Mladosich, Christophe Croux
- Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, and Detection, from Pragmatics to AI Agents Babak Abbaschian
- Algorithm Design and Physician Liability Shujie Luan, Shubhranshu Singh, Tinglong Dai
- Robust Dual-Model Collaborative Random Vector Functional Link Network A. Quadir, A. Rahaman, Mushir Akhtar et al.
- Exploring ESC Winners with Nested Diagrams Anurag Sharma, Marcel N\"ohre, Gerd Stumme
- SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers Kiran Nair, Rodrigue Rizk, KC Santosh
- StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition Zefang Liu, Chenyang Zhu, Sangwoo Cho et al.
- GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis Haochen Zhang, Gengwei Zhang, Laura Yao et al.
- FLARE MCMC: Fidelity-based Layer-Adaptive REcursive proposals for MCMC Harini Venkatesan, Christian Shelton, Ming-Feng Ho et al.
- When Lexical Change Misleads: Rethinking Dynamic Topic Model Evaluation with Traditional and LLM-Based Metrics Charu Karakkaparambil James
- Joint Optimization of Memory and Computing Frequency for Energy-Efficient DNN Inference Yunchu Han, Zhaojun Nan, Sheng Zhou et al.
- Polar Code Based Federated Learning: Convergence Analysis and Resource Allocation Han Xiao, Wei Kang, Nan Liu
- Rewrite Once, Validate Anywhere: Producing OWL-Aware SHACL Constraints (Extended Version) Anouk Oudshoorn, Piotr Gorczyca, D\"orthe Arndt
- Overcoming Shortcut Learning in Graph Neural Networks through Active Explanation Guidance Taraneh Younesian, Steve Azzolin, Antonio Longa et al.
- Revisiting Energy-based Tabular Anomaly Detection: Energy and Reconstruction are Complementary Junichiro Niimi
- Adaptive Protection for Evolutionary Feature Construction in Symbolic Regression with Application to Credit Classification Hengzhe Zhang, Qi Chen, Bing Xue et al.
- Body size predicts how long ant workers live - but not how they age or how they die from heat Alana Moscardi, Rafael da Silva, Gleycon Silva
- Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models Brett Reynolds
- Designing Sustainable Federated Learning as a Service using Neural Architecture Search Keya Patel, Sajib Mistry, Sheik Fattah et al.
- Designing Compact Neural Architectures via Neuron Gating and Mixed Activation Abhishek Shukla, Ankur Sinha, Faiz Hamid
- LP-NAS: Linear Programming-based Neural Architecture Search Abhishek Shukla, Ankur Sinha, Faiz Hamid
- RecipeNet: A Hierarchical Transformer for Recipe Data Pin-Yen Huang, Sachin Chhabra, Prasanth Sai Gouripeddi et al.
Theory 20
Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning
A hidden Markov model (HMM) does three things — infer a belief over hidden state, propagate it through a transition, and emit back into observation space — and the argument here is that time-indexed Predictive Information Bottleneck VJEPA (PIB-VJEPA) has exactly that structure, with the stochastic context encoder acting as an amortized filter, the probabilistic predictor as latent dynamics, and the decoder or inverse target encoder as the emission. Four progressively stronger levels of correspondence are defined along with sufficient conditions for exact sequence-level HMM equivalence, and MCJEPA makes the link concrete by replacing the latent predictor with a learned transition matrix whose powers guarantee exact multi-horizon Chapman-Kolmogorov consistency in the finite time-homogeneous case. Deterministic temporal JEPA falls out as a degenerate Dirac-kernel case, and controlled experiments confirm transition composition, the filtering interpretation, and predictive Markovization on a synthetic process with known structure.
Consistent Model Chasing Is Minimax Optimal: The Exact Value of Scalar Adversarial Adaptive Control under Large Parametric Uncertainty
For the scalar system x_{t+1} = a x_t + u_t + w_t with bounded disturbance and an unknown pole of unknown sign and arbitrarily large magnitude Δ, prior adaptive control theory offered stability certificates, gain bounds, and regret rates but never the exact minimax peak deviation. The analysis pins that game value at **γ*(Δ) = 1 + Δ**, splitting it into the irreducible cost of the disturbance plus the unavoidable price of a single identification spike, and shows the optimal policy is certainty-equivalent deadbeat control at the midpoint of the set-membership consistent interval. Writing the controller as an oracle-selector composition shows this architecture is forced rather than merely sufficient, and the optimal law contains no exploration at all: probing is punished before it pays, commitment is fatal under weak excitation, and optimism costs asymptotically at least twice the optimum.
On the Brittleness of Maximum Likelihood Estimation for Gaussian Process Hyperparameter Optimization
Maximum likelihood estimation (MLE) is the default way to fit Gaussian process (GP) hyperparameters in engineering design, and GPs are widely assumed to resist overfitting, but MLE degrades badly when its assumptions fail. Theoretically grounded alternatives to the MLE objective are compared against it for probabilistic regression and classification, along with practical recipes for choosing among them. In downstream Bayesian optimization tasks the resulting GPs beat tabular foundation models such as TabPFN on prediction accuracy, uncertainty quantification, and inference cost, giving practitioners a concrete blueprint rather than a default loss.
What preferences can - and cannot - predict in multi-agent online learning
Solution concepts in game theory are ordinal — defined by which outcomes players prefer — while no-regret learning dynamics such as follow-the-regularized-leader are continuous-time processes, and it is unclear how much the first determines the second. One direction is settled cleanly: the pure profiles contained in any dynamically stable set must be closed under profitable deviations, and for subgames obtained by restricting action sets, preferences fully characterize asymptotic stability. The converse fails in general — a three-player game is constructed with a preferentially stable set whose span is dynamically unstable — and the gap is bridged by resilience under aggregate deviations, an easy-to-check payoff condition sufficient for asymptotic stability of arbitrary spans of pure strategies.
When Does More Correct Data Hurt? Insertion-Stability and the Limits of Dimension-Based Theory
Adding more correctly labeled data ought to be harmless, but under a monotone adversary that reads an i.i.d. sample and appends any examples the target hypothesis labels correctly, classes of VC dimension at least 2 provably cost a logarithmic factor above the clean PAC rate. Since that bound is a worst case, this work asks which classes actually pay, and answers that it depends on the learner: a learner is called insertion-stable if more correct examples can only shrink its error region, and such learners are immune because risk after insertions never exceeds risk on the clean part alone. Because the Closure algorithm is insertion-stable, every intersection-closed class keeps its clean rate, while classical dimensions cannot predict immunity — two classes can share VC and Littlestone dimension 2 yet split between the clean and penalized rates, and intervals have unbounded Littlestone dimension but are immune; on the known hard class, no monotone permutation-invariant compression scheme of any finite size attains the clean rate.
Boosting Data Augmentation with Stochastic Weight Averaging
Data augmentation is the usual way to bake task symmetries into an ordinary network, and recent theory shows infinitely large deep ensembles trained on augmented data become perfectly symmetric — but ensembles require repeating training many times. Stochastic weight averaging is studied as a cheaper substitute that needs only one run, analysed by approximating the late-training stochastic trajectory with an Ornstein–Uhlenbeck process. In the infinite-width limit, averaging weights over augmented training yields an equivariance gain larger than what the accompanying accuracy improvement alone would predict, a claim backed by experiments spanning computer vision and graph classification under both discrete and continuous symmetries.
AI-Assisted Discovery and Construction of a Counterexample to the Convergence of Three-Block ADMM with the Identity Matrix as its Third Constraint Block
The alternating direction method of multipliers (ADMM) converges for two blocks but can diverge with three, and one subclass had stayed open: whether divergence is still possible when the third constraint block is the identity matrix. Using Codex driving GPT-5.6 Sol, the authors construct an explicit rational instance with strongly convex quadratic first two blocks and verify it along a piecewise-affine reduction path, showing by exact arithmetic that direct three-block ADMM enters a bounded nonconvergent orbit of period 66. A follow-up study of multiplier relaxation shows a problem-dependent small dual step can restore convergence on a fixed instance while no positive relative step works uniformly over the class, and an independent run with Kimi Code on Kimi K3 produced a different certificate (a locally attracting period-23 orbit), suggesting the research harness itself shapes which mathematical objects get found.
LLMs Don't Pay for the Jump
Responding to an argument that language models cannot make the abductive leap that produced Einstein's equivalence principle because they lack embodied simulation, this position piece proposes a different missing ingredient using Planck's 1900 quantization of blackbody radiation as the test case. Planck's postulate needed no sensorimotor grounding — it was forced by a mathematical consequence of classical theory, an infinite predicted energy for a finite measured quantity — and the authors argue neither induction nor deduction could have produced it. They formalize the requirement as a thermodynamic coupling in which epistemic error carries physical cost, and argue fixed-weight transformer inference has no such coupling at any scale, consistent with the empirical observation that output entropy barely changes as accuracy on increasingly hard causal tasks falls from 100% to 17%.
The Dynamics of Intelligence Explosions
If AI systems increasingly do AI research and development, the resulting feedback loop might in principle produce runaway capability growth, and this analysis works through the mathematics of the most explosive cases to see what actually drives them. Singular growth toward a vertical asymptote turns out to be harder to reach than recent economics-inspired models suggest, and there is a neglected middle class of trajectories that grow faster than exponentially yet never hit an asymptote. The pivotal and largely overlooked parameter is generation time, the time to go once around the feedback loop: singular growth is impossible unless generation time rapidly approaches zero.
11 more specialized papers
- Variation Brownian Kernel Ladders Mahdi Mohammadigohari
- High-dimensional nonparametric changepoint detection via low-rank degree-two density projection Guoqing Zhang, Zhaixin Chen
- Identifiability and Order-Dimension Limits of In-Context Learning on Partial Orders Faizanuddin Ansari, Debanjan Dutta, Swagatam Das
- Resource-Adaptive Primal-Dual Learning for One-Warehouse Multi-Store Systems with Censored Demand Jiameng Lyu
- Sequence prediction under a lying oracle Puspabeethi Samanta, Nikhil Karamchandani, Jayakrishnan Nair
- Classical Limits of Spectral Filtering in Quantum Generative Models Marco Roth
- Connected Subspace Clustering: Hardness, a Scalable Heuristic, and an Application to Sea Level Geodesy Johanna Hillebrand, Jan H\"ockendorff, J\"urgen Kusche et al.
- A Generalized Parallelogram Rule for Proportional Analogies on Riemannian Manifolds Pierre-Alexandre Murena, Marcelo Hartmann
- Convex losses and their applications to SVM, SVR, and Shallow Neural Networks Filippo Portera
- Quantum Multi-Armed Bandits and Linear Bandits: Lower Bounds and Algorithms Maoli Liu, Zhuohua Li, John C. S. Lui
- Non-Shattering at and Above the Dynamical Temperature in the Spherical Pure p-Spin Model Taegyun Kim
Reinforcement Learning 16
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
Group-based reinforcement learning assumes the rollouts compared within a group are doing the same kind of thing, but in open-ended interaction an agent might answer immediately, ask a clarifying question, post a progress update, or confirm before acting — all valid, and a reward model's stylistic preferences then contaminate the relative advantages. Framing this as a reward fairness problem, ARC (Advantage Regularization via Conditioning) groups rollouts by interaction strategy before computing advantages, combined with hybrid rewards and entropy regularization, and is studied inside an interaction paradigm that separates what the user sees from latent reasoning and tool calls, supported by an 86K-example strategy-annotated training corpus. ARC substantially improves scores on the tau and tau^2 tool-use benchmarks, while the decoupled interaction design cuts time-to-first-token from 4.91 to 1.27 seconds versus a think-first baseline.
Reward Machines for Signal Temporal Logic
Signal temporal logic (STL) specifies real-time properties over real-valued signals and supplies a quantitative robustness score, which prior work has fed directly to reinforcement learning as a reward; the trouble is that robustness depends on the entire execution history, so the state space blows up for long-horizon specifications with nested temporal operators. The proposed approach compiles a specification into a timed alternating automaton, augments the state space with automaton locations and clock valuations to serve as a compact memory, and derives Markovian rewards from the automaton's acceptance condition. Policies trained this way achieve higher robustness scores and satisfaction rates than those learned from robustness-based rewards.
GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings
Reinforcement Learning with Verifiable Rewards (RLVR), typically optimized with Group Relative Policy Optimization (GRPO), is the standard recipe for improving reasoning in language models, but nearly all published studies train and evaluate in English. A large-scale sweep varies the base model, the training language, and the reward applied to the language the model reasons in. Training a model to reason in its native language costs only a small amount relative to reasoning in English, and training in a single language often transfers gains to many other languages. The effects are strongly model- and language-specific, though, with some training languages causing severe regressions on out-of-domain capabilities elsewhere, so multilingual RLVR needs broad evaluation to catch them.
Dynamic Multi-Depot Vehicle Routing with Online Requests: Event-Driven Transformer--DRL and Rolling-Horizon Benchmarking
Dispatching vehicles from multiple depots when requests arrive over time requires decisions in milliseconds, making it a natural target for learned policies. Masked multilayer perceptron and Transformer policies were trained by behavior cloning and proximal policy optimization (PPO) with deterministic feasibility masking and fixed-prefix route commitments, then compared against insertion heuristics and time-limited rolling-horizon optimization on a shared 20-scenario protocol. Every method served all requests without invalid actions, but a simple nearest-feasible heuristic achieved the lowest objective and beat both learned policies on routing quality, waiting time, stability, makespan, and runtime; rolling-horizon optimization won on waiting time and makespan at far higher compute. The learned policies did keep millisecond inference and transferred to 80-request instances without retraining, and PPO helped the Transformer on average while adding seed variance.
Learning to Run Power Networks: Effective AlphaZero-inspired Topological Control
Reconfiguring grid topology is a cheaper way to relieve congestion than redispatching generation, but the action space is combinatorially huge and operational constraints are strict. This study applies AlphaZero-style planning with Monte Carlo Tree Search (MCTS) to proactive grid operation and systematically varies reward design, observation density, and search guidance. The tuned agent reaches 98.43% peak survivability, well above a proximal policy optimization (PPO) baseline, and — counterintuitively — running MCTS without a learned prior policy or value function improved training efficiency, while a plain binary survival reward guided search better than multi-objective alternatives. The authors conclude that pure reinforcement learning is insufficient and that domain heuristics, binary rewards, and a restricted line-load observation space are what make the system work.
Reinforcement Learning-Based Production Scheduling in an Industry-Based Coating Scenario Using the Digital Model Playground
Reinforcement learning results for production scheduling mostly come from simplified benchmark shop problems, which limits what factories can conclude from them. This work models an industry-inspired coating process with sequence-dependent setup times, machine breakdowns, due dates, and variable utilization in the open-source Digital Model Playground (DMPG) discrete-event simulation framework, then trains Deep Q-Networks and Proximal Policy Optimization agents against conventional dispatching rules. Both agents improve key performance indicators in a balanced way, with PPO giving the most robust performance, and the shareable scenario plus framework is offered as a reusable testbed.
Traj-LeWM: Path-Aware World-Model Planning via Latent Trajectory Cost
LeWM is a pixel-based latent world model that scores candidate action sequences purely by how close their predicted endpoint lands to the goal, so it learns only local next-step transitions and ignores what the rest of the predicted path looks like — a problem when predictions diverge from execution. Traj-LeWM keeps that objective and endpoint score but adds a goal-conditioned latent trajectory cost that summarizes the whole predicted rollout, used during training as trajectory-preference supervision and during planning as a second ranking signal. It beats LeWM by 3, 14, 7, and 7 percentage points on Push-T, OGBench-Cube, Reacher, and Two-Room, with ablations separating the representation-shaping and candidate-ranking contributions.
Deep Reinforcement Learning solution for pickup and delivery routing problems with time window and capacity constraints
Real-time vehicle routing for goods pickup and delivery becomes intractable for classical methods once capacity and time-window constraints are added at medium to large scale. A modified version of the JAMPR deep reinforcement learning model is applied to the Pickup and Delivery Problem with Capacity and Time Window constraints (CPDPTW), reportedly the first successful deep RL solution for this constraint combination. The learned policy produces fast optimal solutions on small and medium instances and fast suboptimal ones beyond roughly 200 nodes.
Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine
Reinforcement learning for mechanical ventilation settings usually consumes only structured electronic health record data, discarding the clinical context that lives in free-text notes; naively adding those notes fails because copy-forward text, templating, and repeated documentation swamp the genuinely new information at each timestep. The proposed framework strips temporal note redundancy before policy learning, comparing an embedding-space decomposition using singular value decomposition over local history subspaces against an interpretable sentence-level diff that filters previously documented sentences prior to encoding. On real intensive care unit data, the redundancy-stripped state representations beat both structured-only and raw-note baselines across all four off-policy evaluation methods tested (model-based rollouts, fitted Q-evaluation, weighted importance sampling, and weighted doubly robust evaluation).
Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL
Training terminal-using agents with reinforcement learning requires executable environments whose rewards are trustworthy and whose difficulty sits where the current policy actually learns, but recipes like Self-Instruct and Evol-Instruct apply one prompting strategy to every seed regardless of policy state. Envs-FORGE turns verifier rewards into per-seed synthesis decisions: it estimates each seed's pass rate, scores six directions relative to a target learning frontier, and solves a small mixed-integer linear program to pick the action that conditions generation, which then rewrites the instruction, fixtures, oracle solution, tests, and Docker image together so only gold-verified bundles enter training. On Qwen 3.5 35B it lifts Pass@1 by 9.2 points on tb-core and 6.4 on tb-2.0, beating the best fixed recipe by roughly two points, and reaches 77.1% on SWE-bench Verified against 73.4% for the base model. All compared methods export 100 verified environments at comparable token cost, holding training-set size fixed.
CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving
Long-horizon goal-directed driving asks a reinforcement learning policy to learn several competing behaviours simultaneously — reaching a distant goal, following a route, avoiding obstacles, obeying traffic signals — and a fixed reward gives no ordering over them. CORAL advances two schedules in lockstep: a five-stage curriculum that lengthens routes and tightens behavioural constraints, and a stage-aware reward whose weights shift from mission progress toward route following, safety, smoothness, and rule compliance as difficulty rises. The policy is a multi-stream actor-critic trained with Proximal Policy Optimization in CARLA on a compact 99-dimensional state — a polar LiDAR histogram plus telemetry, route geometry, and rule indicators, with no point-cloud encoder or bird's-eye-view raster. It succeeds in all twenty evaluation episodes on the hardest routes where two PPO baselines reach 5% and 10%, drops to 55% with both schedules disabled, and transfers zero-shot from its training town to seven unseen towns at 68–98% success.
Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View
Reinforcement learning post-training for diffusion models has split into two apparently unrelated camps — reverse-trajectory methods based on discretized likelihood ratios, and forward-matching methods trained on reward-labeled noised versions of rollout samples. Starting from the regularized diffusion-RL objective and applying importance sampling between sampling stochastic differential equations yields an explicit policy-gradient estimator on trajectory space that contains the Itô integral behind Flow-GRPO-style updates, and an equivalent variance-reduced value-gradient form that reproduces the forward-matching structure of AWM and DiffusionNFT. The empirical gap between the two families is therefore a variance-reduction effect, not a difference in underlying RL principle. The resulting design space, organized by value-gradient estimation, weight function, and sampling choice, yields a multi-sample kernel density estimation value-gradient estimator with scale-bounded weights that improves on prior baselines on SD3.5-M and Qwen-Image.
Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training
On-policy reinforcement learning runtimes execute rollout, reference scoring, and actor training as strictly serial phases, which wastes capacity for vision-language models where processing dense video inputs and prompt prefixes consumes much of each phase. Rollplex observes that prefix computation does not depend on the generated response, so it can overlap with rollout decoding without breaking synchronous on-policy semantics; making that schedule fit memory requires phase-aware control of high-bandwidth memory residency plus parallelism-aware weight sharing that reuses physical storage for tensors whose layouts are compatible across different tensor-parallel degrees, rebuilding only the incompatible ones. Colocating Qwen2.5-VL-32B naively would need roughly 165 GiB per GPU; on 32 H800 GPUs the runtime instead delivers 1.23×–1.30× speedup over serial colocation and 1.57×–2.24× over disaggregation at the same GPU budget.
3 more specialized papers
- Continual Evolution Strategies in Control Tasks Nicola Pitzalis, Eleni Nisioti, Antonio Carta et al.
- Offline Deep Q* Estimation with Diffusion Models Xiaohong Chen, Yuling Jiao, Lican Kang et al.
- Online Inference in Distributional Temporal-Difference Learning Yang Peng, Liangyu Zhang
Safety & Alignment 16
Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation
Large Audio Language Models (LALMs) are increasingly used for speech recognition and audio question answering, but checking whether they serve demographic subgroups equally is confounded by what is actually said and by speaker-specific vocal traits. The proposed evaluation framework is a semantic-aware mixed-effects regression that adds sentence-level semantic embeddings of the reference text as covariates and treats speaker identity as a random effect, with the embeddings drawn from the same LALM being evaluated so that semantic variation is controlled as that model perceives it. Across simulated data and real-world benchmarks, the method substantially reduces spurious fairness findings and produces subgroup performance gaps that are more stable and easier to interpret.
Language-Specific Gaps in AI Safety Training Datasets
Model providers cite multilingual safety benchmarks covering a dozen or more languages as evidence of safety for non-English speakers, but collection-level coverage claims can hide weakness in any individual language. An audit of 21 resources across 25 language slices, spanning Hausa (low-resource), Swahili (mid) and French (high), inspects provenance, annotation reliability, access, harm-taxonomy coverage and data reuse one language at a time. Gaps only partly track resource tier: within a single pipeline the Hausa slice fell below its own paper's translation-quality acceptance threshold while the Swahili output cleared it comfortably, and self-harm and sexual-content categories had no native-language coverage in either African-language tier. The authors argue this thinness lines up with the known persistence of multi-turn jailbreaks in non-English languages, and release a reusable slice-level audit protocol plus the safety-slice-audit dataset.
Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation
The EU AI Act obliges providers of high-risk systems to document how their systems reach decisions, and circuit discovery in mechanistic interpretability is the natural source of such evidence — provided two competent analysts using the same tool with different defensible settings would file the same thing. A pre-registered grid crossed seven analytic axes, each level drawn from a published implementation, over GPT-2 small on the indirect object identification task, mapping every discovered circuit through a deterministic claim map into a structured Annex IV statement. Across 15,840 specifications, 7,561 of which produced a claim, the derived statement flips across 73.2% of specification pairs and the most common claim covers only 41.1% of the space; standardizing the evaluation metric leaves the flip rate at 59.4%, and dropping circuit size from the claim entirely still leaves 27.1%. The underlying circuits are near-disjoint (median pairwise Jaccard overlap 4%) and functionally uncorrelated (Cohen's kappa 0.015), so this is not one mechanism described in different words; the study covers one model and one task.
ASSERT: A Measurement Pipeline for GenAI Audits
Audits of generative AI systems typically report a single compliance rate that gets used to compare systems, track regressions, and gate deployment, even though that rate reflects the auditor's measurement choices as much as the system itself. ASSERT is a specification-driven pipeline that helps draft a behavioral rubric and test cases, runs the audit, and binds every reported rate to a written record of the choices that produced it. In a case study on conversational deception, varying the dialogue setup, the simulated user, the judge, or the evidence bar for non-compliance shifts the reported rate enough to reorder which systems look better, which is exactly what the attached specification makes attributable.
Mandato: Protocol-Level Enforcement of Digitally Signed Mandates on AI Agent Actions with Cryptographically Chained Audit Trails
Agents now act on real systems through tool-calling protocols such as the Model Context Protocol (MCP), yet whether a given call was actually authorized lives in application code that is unsigned, unauditable, and logged without evidentiary weight. Mandato is a transparent MCP proxy that enforces digitally signed mandates — machine-readable artifacts stating which tools an agent may call, under what parameter and context constraints, for how long, and on whose behalf — evaluating each call against the mandate chain, blocking non-conforming ones inline, and writing every permit and deny decision into an append-only hash-chained audit log anchored with qualified timestamps. The mandate model is deliberately shaped after the civil-law delegation of authority so lawyers and auditors can read it, and the work maps the mechanism onto EU AI Act Articles 12 and 14, GDPR accountability, NIS2, and eIDAS 2; the reference implementation is described alongside a planned quantitative evaluation of overhead and tamper-evidence cost rather than measured results.
Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers
Safety classifiers shipped alongside large language models tend to enforce the policy they were trained on rather than the one a deployer wants, and they decay as traffic shifts. Regime-Conditional Verification (RCV) wraps an off-the-shelf classifier without retraining it: from the classifier's internal representations it estimates the probability that each prediction disagrees with the deployer's policy, selectively corrects the likely-wrong ones, and reuses the same estimates as a label-free distribution-shift detector that triggers fine-tuning only when cheaper repairs fail. Across three classifiers and two benchmark datasets it improved policy adherence in every combination, catching up to 0.81 of previously missed unsafe content with the underlying classifier untouched, and it flagged all ten held-out attack campaigns in a deployment study.
BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs
Work on social bias in language models mostly inspects final answers, leaving open which internal reasoning steps actually generate a biased outcome. BiasTrace is an annotation scheme that labels behaviours inside model-generated reasoning traces — both bias-specific ones such as unsupported demographic assumptions and general patterns such as overthinking — and links them to biased outputs, scaled up with validated LLM-as-a-judge labelling to build a large annotated dataset. The analysis finds that biased outputs typically arise from subtle reasoning behaviours rather than explicitly biased language, that reasoning-level labels improve bias detection, and that the annotated behaviours can be targeted for inference-time mitigation.
Training Fair Tabular Foundation Models
Tabular foundation models (TFMs) now lead on tabular prediction via in-context learning and are being used for high-stakes decisions, yet their fairness behaviour is largely uncharacterized. FairTFM builds fairness into TFM training itself so predictions come out fair in a single forward pass, working around two obstacles: sensitive attributes are rarely available in training data, and standard fairness techniques assume task-specific training rather than in-context learning. The recipe combines synthetic fairness tasks with a gradient reversal layer that pushes the model toward representations invariant to sensitive attributes, and across 132 fairness tasks it improves fairness consistently while keeping accuracy competitive.
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
Facts seen frequently during pretraining are memorized more deeply and resist removal longer, yet unlearning methods apply the same gradient pressure to every target regardless of how common it was in training data. AdaPop combines local token confidence with a per-fact exponent scaled by an external popularity proxy such as Wikidata sitelinks or an LLM judge, and uses a dual-ascent controller to retune the retain penalty each epoch instead of hand-tuning the forget-retain tradeoff. Across three model families and two benchmarks it leaks roughly 5x less forgotten content under paraphrased queries and 1.6x less under adversarial reformulations than competing methods, with hidden-state analysis showing forget-set representations move further from the original model while retain-set representations stay in place.
Detecting Contaminated Code-Generation Prompt Batches via Influence Functions
Prompts can steer code-generating language models toward insecure implementations, and defenses built around known vulnerability patterns or a fixed threat model miss attacks they were not designed for. CodeSIFT avoids specifying vulnerabilities at all: it uses influence functions to measure the parameter-space influence of code a model generates, then applies a statistical test to decide whether a candidate batch of prompts deviates from a benign reference distribution. Across three open-weight code models from 3B to 7B parameters and two new benchmark datasets covering varied vulnerability types, it reached AUROC up to 0.98 at moderate-to-high injection rates with well-calibrated false positive rates, substantially beating static analysis baselines.
A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation
Regulatory standards phrased as principles — "fair, clear, and not misleading", "deliver good outcomes" — resist encoding as binary rules, so language models are increasingly used as the judge, yet they are rarely tested for anything beyond accuracy. Principle-Bench supplies 168 cryptoasset financial-promotion scenarios mapped to two UK Financial Conduct Authority principles, with paraphrase, adversarial keyword-stuffing, and boundary perturbations written to a pre-registered rubric, and evaluates four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. Alongside it, Ceca (Calibrated Exemplar-Cluster Assessment) provides a calibrated assessor with per-exemplar counterfactual attributions. No method wins on all four axes, and a 120B model judge, the strongest on benign inputs, drops from 0.74 to 0.27 accuracy on keyword-stuffed Consumer Duty prompts — behaviour the authors call "compliance theatre" — while a judge from another model family agrees with it at only Cohen's kappa 0.16 on that split, localising the failure to the model rather than the corpus.
Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons
Neuron-level jailbreak defenses promise surgical precision but in practice either suppress toxic semantics spread across many pathways (requiring a large intervention footprint) or rely on external classifiers that also damage utility-critical neurons, and both stay switched on for benign traffic. Tripwire instead runs per-neuron hypothesis tests with false-discovery-rate control plus a utility-specificity filter to isolate genuinely safety-specific neurons, then clamps them to their harmful-conditional mean activations to inject an internal "this input is harmful" signal that fires the model's own aligned refusal. The clamp ships as two provably equivalent modes: a detector-gated inference-time intervention and an offline bias-patch weight edit. Across four safety-aligned models and four attacks it is training-free and cuts average attack success rate to at most 2.0% while costing only 0.5% to 5.3% on MT-Bench.
Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice
As patients start asking chat assistants which doctor to see, those assistants become infomediaries that silently decide which physicians are visible, so the authors run a prespecified randomized audit of what actually drives the choice. Seven models (six open-weight plus gpt-4o-mini) picked among five synthetic family-medicine physician cards with independently randomized attributes across 3,024 choice sets, three patient personas, and nine paraphrases, producing 40,068 scored responses, with gender and ethnicity signaled through names as in correspondence audits. Reputation dominates — a rating rise from 3.9 to 4.7 adds 31.4 percentage points of choice probability and a fee rise from $90 to $190 subtracts 20.0 — but demographic parity still fails, with female-signaled and Hispanic-, South-Asian- and Black-signaled names gaining 1.3 to 2.9 points over White-signaled ones, and mere first-listed position worth about $11 in fee-equivalent terms. The models cited gender or ethnicity in at most 0.03% of their stated reasons, so transparency regimes built on model self-report would miss these tilts entirely.
Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers
Moral preference elicitation — polling people on hypothetical dilemmas and training a policy on the aggregated votes — is often treated as a neutral way to align AI with public values, but three upstream developer choices shape the result: which features go to a vote, which voters are sampled, and how the question is worded. Across two phases with 809 participants and three deployment contexts (AI kidney allocation, AI agents standing in for absent workers, and generative depictions of the deceased), the authors measure how each stage moves the outcome. Morally relevant features did not transfer across contexts, preferences split by political ideology for roughly a third of features with some differences reversing sign, and question framing alone widened or narrowed ideological gaps by up to a full scale point. The conclusion is that voting-based alignment cannot deliver fairness by aggregation alone, and that each pipeline stage should be audited and disclosed.
2 more specialized papers
- CutClean: Neural Network Pruning for Privacy-Preserving Inference Leonardo Magliolo, Vito Paolo Pastore, Giuseppe Valenzise et al.
- Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence James K. Wiles
Vision 16
TRUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection
Computer-aided detection systems for colonoscopy are usually trained and scored on curated lesion-centric clips, which omit the long stretches of healthy tissue and the procedural artifacts that dominate a real examination. TRUE-Colon is a benchmarking protocol that measures deployment-relevant behavior alongside localization accuracy, applied to Faster R-CNN, YOLOv8, YOLOv11 and RT-DETR across the curated SUN and PICCOLO sets and 60 unedited full-length procedures from REAL-Colon. Models trained only on curated clips collapse on full procedures, while procedure-trained models reject non-polyp content far better and still retain accuracy on the curated benchmarks — an asymmetry consistent across architectures. The Transformer detector gave the strongest sensitivity and earliest, most persistent detections, with the convolutional detectors competitive at higher throughput.
SDO: Subspace Deconflicting Operator for Multi-Adapter Composition
Loading several independently trained adapters into one diffusion backbone is an appealing route to multi-character generation, but deploying them together causes identity mixing, attribute leakage between characters, and unstable scenes. Framing the interference in parameter space, SDO reconstructs each adapter's layer-wise low-rank update, extracts compact subspace signatures, scores pairwise conflict by output-subspace overlap, and applies a permutation-equivariant transformation that suppresses harmful shared directions while preserving identity-specific ones, mapping results back into ordinary adapter weights usable in existing inference pipelines. Experiments report improved identity fidelity and compositional stability, with the gap over naive joint deployment widening as more adapters are composed.
Post-training Quantization for Hybrid Iterative Generative Models
Hybrid generative image models that couple autoregressive and diffusion paradigms produce high fidelity but pay for it in iterative inference cost, and applying standard post-training quantization to them causes outright model collapse. Diagnosing the failures identifies two causes: excessive activation outliers that force an unwinnable trade-off between covering the outliers and preserving normal-range precision, and amplified anomalies where small quantization errors compound into a calibration-inference mismatch. HyGenQ counters these with hierarchical cluster decoupling, which isolates outlier channels through multi-stage clustering, and scaling recalibration, which rescales anomalies past the Gaussian bound instead of truncating them, successfully quantizing representative hybrid models to 8-bit weights and activations while outperforming existing baselines across model families.
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
Interactive video world models need to generate frames causally with very low latency while still responding correctly to keyboard and mouse input, which is hard to preserve when distilling a bidirectional generator into a one- or few-step sampler. ForgeWM is a four-stage recipe — domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching against a bidirectional teacher — producing students specialized for denoising budgets of 1, 2, and 4 steps, plus a dual-path deployment mode where the one-step draft is re-noised and refined at replay time. On paired Minecraft trajectories the students lead the compared systems on imaging quality, motion-profile agreement, action-sign accuracy, and mouse-control accuracy while reaching the lowest LPIPS, and replay-time refinement matches four-step quality while staying about three times closer to the trajectory the player actually experienced than regenerating from noise. The same recipe transfers to gamepad-controlled first-person shooter gameplay.
Adversarial Learning of Classifier-Free Guidance Schedules
Text-to-image diffusion models use classifier-free guidance, or CFG, with a single fixed scale applied across every timestep, sample, and prompt, which is rarely optimal and can produce artifacts. The guidance scale is instead learned as a function of diffusion time, conditioning, and the current noisy sample by casting the problem as density-ratio estimation: a discriminator estimates the time-dependent log-density ratio between the true and guided marginals, while a small generator network predicts the state-dependent scale. The learned schedules beat both hand-designed CFG heuristics and prior dynamic-guidance methods on standard text-to-image benchmarks.
AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations
Metrics used to judge face editing and privacy-protection systems are meant to stand in for human perception, but current representation-based ones model behavior without cognitive structure and assume a single universal observer, so they misrepresent how different populations actually perceive facial similarity. AlignFace builds three findings from the cognitive psychology of face perception — dependence on featural and configural attributes, nonlinear psychophysical scaling, and own-group bias — directly into its architecture, pairing a vision-language encoder and gated cross-attention with a concept bottleneck over interpretable face attributes and a neural generalized additive model for their nonlinear contributions, and is trained on a new FACETS dataset. Experiments report significantly better alignment with the perceptions of human subpopulations than baseline and recent learned perceptual metrics, while keeping the reasoning inspectable.
QuaSAR: Quantization Compensation via Stable Activation-Aware Rank Truncation
Training-free post-training quantization methods often use a goodness-of-fit score to decide which layers get closed-form residual compensation, but at 4-bit weights and activations (W4A4) that gate conflates genuinely unpredictable quantization error with numerical breakdown of the solver. Rank-deficient input activations produce ill-conditioned or singular Gram matrices, yielding spuriously negative fit scores so that layers which could in fact be compensated are thrown away. QuaSAR replaces the solver with a parameter-free truncated pseudoinverse that drops collapsed directions before inversion, reaching 81.42% top-1 accuracy on ViT-B under W4A4 and beating both prior post-training and fine-tuning-based baselines; combined with joint low-rank compression it reaches 80.26% at 54.7 MB.
Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation
Text-to-image diffusion models offer no dial for continuous concept-specific control such as aesthetic quality, and they remain unreliable on structures needing local coherence like rendered text or hands. Using a new notion of concept-wise mutual information, the authors find that generation of specific structures is localized to distinct layers, then build Concept Guidance (CoG): quantify each layer's concept-specific impact, and steer denoising with a weighted combination of predictions produced while concept-relevant layers are skipped. The method needs no training, gradients, external models, or prompt engineering and improves several targets on PixArt-alpha, SD3, SD3.5, and FLUX.1-dev out of the box.
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Modern video generators can produce convincing footage of wars, disasters, and other emergencies, but existing detection benchmarks say little about how detectors behave on this kind of content or after it spreads online. RA-Bench pairs 1,830 real videos across 10 social-risk categories with 16,056 clips generated from four open-source and five closed-source generators, then evaluates seven traditional detectors, ten zero-shot multimodal models, and two multimodal large language models fine-tuned for the task. None of the three detector families generalizes consistently across the benchmark, generation quality and conditioning affect each family differently while source-level patterns stay stable across sampling seeds, and the videos that fool human judges are also the hardest for detectors — with social dissemination (re-encoding and sharing) making detection harder still.
Style or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision Embeddings
Frozen image embeddings from encoders like CLIP report high accuracy on classifying paintings by art-historical movement, but the usual random train/test splits place works by the same painter on both sides, letting a classifier win by recognising the artist instead of the style. Re-running the task under an artist-disjoint protocol — holding out every artist in turn across a balanced set of 320 paintings from four twentieth-century movements — drops 5-nearest-neighbour accuracy from 0.87 to 0.77, with the loss concentrated almost entirely in Surrealism (down twenty points) while Impressionism and Cubism hold steady. The pattern repeats across four encoders including a vision-only self-supervised model, which locates the effect in visual structure rather than language supervision.
Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Interactive game world models usually autoregress pixels or latents directly, which forces pose, geometry, and occlusion to be tracked implicitly and lets errors compound over long horizons. Marionette splits the job three ways: a two-stage autoregressive dynamics model predicts an explicit 276-dimensional 3D world state of articulated skeletons, metric root trajectories, and rotations; a zero-parameter graphics bridge turns that state into pose-control videos with world-space geometry and occlusion computed in closed form; and a control-conditioned video diffusion model paints photorealistic frames. Because behavior lives in the explicit state, it can be corrected there — left unconstrained, two generated characters drifted 21.2 m apart against roughly 5 m in recordings with a third of frames showing ground penetration, while adding just a terrain collider and a separation cap to the state cut penetration by 66% with no change to the observation model. Forcing a mismatched action stream shifted root-aligned joint error by 31% across 48 held-out segments, and routing appearance through the predicted state cost little visual fidelity (FVD 831 versus 799 for recorded poses).
5 more specialized papers
- UltraArUco: A Lightweight Multilingual Library And Framework With Low-Latency Real-Time Marker-Based Tracking System For Mobile AR Interaction Mikhail Kiselev, Aleksandr Marukhin, Ivan Snegirev et al.
- Secret-Stego Dissimilarity as a Design Axis: Invertible Coverless Image Steganography with Diffusion Models Hongxin Xu, Jianping Mei, Can Wang et al.
- What to Preserve, Where to Adapt: A Depth-Wise Analysis of Forgetting in Continual Gynecological Image Segmentation Amal Saqib, Tausifa Jan Saleem, Numan Saeed et al.
- HiCo-GS: Hierarchical Context Aggregation and Geometric Consistency for Octree Gaussian Splatting Wei Zhang, Shengkai Yu, Shiqiang Gong et al.
- GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection Yingjie Ma, Zitong Yu, Wei Jia et al.
Multimodal 13
VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation
Systems that synthesize speech from a written description of a voice currently cover a narrow range of speakers and offer little control after generation, such as cloning a voice or adjusting its emotion. VoiceDesigner addresses both with a hybrid data pipeline that combines digital signal processing with speech generation models to build a dataset spanning real and fictional voices, plus a diffusion transformer modified to handle complex conditioning across generation and editing in one model. Subjective and objective evaluations show better alignment with both voice descriptions and editing instructions than state-of-the-art text-to-voice systems, at comparable perceptual quality and usability.
VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents
Speech language models mostly operate turn by turn and cannot handle a user talking over them, while duplex speech-to-speech systems that fix latency tend to sacrifice audio quality because recognition, interruption handling, and synthesis must all be optimized jointly. VoiceChat-TTS keeps synthesis modular: it consumes an LLM's text-token stream directly, emits silence when no text is available, and accepts explicit control tokens for interruption. The design gives always-on streaming speech at low latency and handles mid-utterance barge-in without resetting the KV cache, so the model does not restart generation when interrupted.
Content Based Video Narration of Gameplay with Vision Language Models
Live esports-style commentary exists for professional broadcasts and almost nowhere else, so this system produces spoken narration for arbitrary gameplay recordings using a general-purpose vision-language model and text-to-speech, with no engine telemetry, game-specific instrumentation, or task-specific training. Three mechanisms carry it: temporal mosaic packing arranges nine sampled frames into one 3x3 image so an image-native model can reason about motion from a single payload, context-conditioned prompting replays the most recent narrations as assistant history to suppress repetition, and duration-conditioned generation with elastic alignment time-scales or pads synthesized audio to fill each segment exactly without a forced aligner. The speech stage can run fully locally via a 6-bit quantized 4B-parameter model on Apple silicon, the mosaic cuts per-minute image payloads by 9x, and the release includes a candid account of failure modes such as hallucinated game state, mosaic resolution loss, and prosody artifacts.
HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation
Multimodal retrieval-augmented generation typically shreds structured documents into loose text and image chunks, discarding the section hierarchy and the local pairing between a figure and the text around it, which hurts both which evidence is picked and where images are placed in the answer. HAM-RAG carries document hierarchy through retrieval and generation as a grounding signal, keeping each piece of evidence tagged with its position and its neighboring text-image relations in the prompt, and ships HAM-Bench, a benchmark spanning game walkthroughs, web pages, scientific papers, and step-by-step recipes. Across several backbones it raises the main multimodal average by 17.3% over the strongest non-hierarchical baseline, with a 24.2% gain in image-text alignment on the game-walkthrough split.
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
Recovering geometry, matching points across views, and reasoning about spatial relations are usually handled by separate task-specific models or bolted-on geometry modules, which blocks any sharing between these complementary views of the same scene. SPARGen recasts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation inside one native multimodal generative model, emitting compact structured and linguistic answers as token sequences while producing dense geometric fields in image-aligned form so that spatial supervision shapes a single shared representation across all three task families. Evaluations across reconstruction, correspondence, and spatial reasoning benchmarks report competitive performance from that unified model.
Self-Supervised Visual On-Policy Distillation
On-policy distillation for visual tasks normally needs an asymmetry between teacher and student — a larger teacher, ground-truth answers, or annotated regions of interest — which is unavailable when no privileged signal exists. Self-Supervised Visual On-Policy Distillation (S$^2$VOPD) inverts the setup by removing information from the student instead of adding it to the teacher: the teacher's distribution conditioned on the original image is distilled on-policy into the student's distribution conditioned on a strongly augmented view of the same image. A sweep over augmentation families shows asymmetry is what matters (symmetric self-distillation hurts), strength peaks at moderate levels, and augmentations that erase the question-relevant evidence produce large but useless discrepancies. Across six fine-grained perception benchmarks it lifts Qwen3.5-4B from 70.7% to 77.4%, beating open-source models up to Qwen3-VL at 235B as well as GPT-5.4, and recovering 96% of the gain from privileged-information methods.
Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding
Millimeter-wave (mmWave) radar sees through darkness and occlusion, but pairing it with language models has been blocked by scarce radar-text data, incompatible dataset conventions, and the absence of a foundational radar encoder. The approach sidesteps all three with a minimal textualization interface that serializes each mmWave point cloud into short natural language so off-the-shelf LLMs can answer questions about it, and packages this as mmWave-QA, the first benchmark for language-conditioned mmWave human perception, built by harmonizing heterogeneous public datasets through calibration-aware preprocessing and a shared taxonomy across six scenarios and five question-answering tasks. Evaluation of current LLMs on the benchmark indicates non-trivial zero-shot reasoning over radar data and robustness where visual sensing degrades.
Seeing Red, Thinking Bad: Color Bias in Vision Language Models
Vision language models are being used for screening and recommendation tasks where text often arrives as an image, raising the question of whether visual styling changes their reading of semantically identical content. Stealth Visual Prompts alter only the color and contrast of rendered words while preserving wording, and applying them systematically showed that coloring positive words green consistently pushes sentiment predictions positive, with models frequently failing to register negative words in the same text. Latent analysis ties the shift to color-induced changes in the vision encoder's representations, and lowering text-background contrast increases reliance on visually salient cues and produces more incorrect visual question answering outputs.
5 more specialized papers
- MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation Rafi Ibn Sultan, Hui Zhu, Chengyin Li et al.
- S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling Xueqi Wang, Zhigang Wang, Runqing Zhang et al.
- A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade et al.
- Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge Kexin Shi, Renhe Sun, Yuge Huang et al.
- Disentangled Shared Representations Improve Morpho-Transcriptomic Integration Julian Ostermaier, Swann Ruyter, Reuben Dorent et al.
Robotics 12
Active Perception for Embodied Disambiguation
When a robot is told to fetch something and cannot tell which object is meant, the missing information is often physical rather than linguistic — the target is occluded, the viewpoint is bad, or a label is unreadable — so asking the user cannot resolve it. The proposed framework makes moving to a better viewpoint the primary information-gathering action, with a vision-language model deciding from accumulated visual evidence and dialogue history whether to keep observing, request clarification, or commit to a target. Real-robot experiments show physical information gathering and user-intent clarification working within one disambiguation loop, with active observation also improving the quality of clarification questions by revealing object names and attributes.
hint$^2$: Hierarchical World Models for Inference-Time Temporal Logic Guidance
Linear Temporal Logic (LTL) can express the temporal structure and safety constraints that language-conditioned manipulation policies handle poorly, but LTL is evaluated over long trajectories while modern policies emit short action chunks and replan in closed loop. hint2 bridges that gap at inference time with two world models at different abstraction levels: a high-level model predicts how candidate actions change task-relevant atomic propositions to drive progress through the LTL automaton, and a low-level dynamics model predicts immediate state evolution for local safety guidance. It outperforms existing LTL-guided diffusion and inference-time steering methods on CALVIN, satisfies instructions with combined liveness and safety constraints, and transfers to a real UR5e arm.
Coverage Aware Active Evaluation for Failure Discovery with Paired Systems
Autonomous systems fail rarely and in varied ways, so finding those failures under a limited real-world testing budget is hard, and cheap proxies such as simulators or lower-fidelity policies surface failures that often do not transfer. The proposed method learns a local predictor of target-system risk by correcting proxy failure signals with control-variate-inspired residual modeling, then selects scenarios using a support-aware mutual-information objective that favors realistic, well-supported regions while spreading coverage across distinct failure modes. On autonomous driving, manipulation and quadruped velocity-tracking tasks it discovers up to twice as many failures as random sampling and active-learning baselines, including severe modes the baselines miss entirely.
AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning
Dexterous manipulation policies are expensive to train because robot teleoperation data is scarce and action spaces differ across hands and grippers, and models trained on mixed sources tend to tie task cues to embodiment-specific appearance. AdvDex is a Vision-Language-Action framework built on three pieces: OmniShare, a large multimodal dataset of human manipulation demonstrations with kinematic and tactile supervision; a Joint-Aligned Action Space of an SE(3) wrist pose plus 15 finger joints that puts human hands, robot hands, and parallel grippers in one representation; and domain-adversarial training that strips embodiment cues from the visual encoder. Experiments on hand-action prediction and real hardware report gains over baselines, zero-shot transfer of skills learned from humans to robots, generalization to unseen objects and scenes, and few-shot adaptation from little data.
Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use
End-to-end Vision-Language-Action models must solve for continuous low-level actions directly, which makes them data-hungry and brittle outside their training distribution. ART (Agentic Robot with Tool-use) is a tool-injection layer that fine-tunes any such model to call off-the-shelf modules for low-level vision, high-level affordance estimation, and embodiment-specific control, shrinking the action space the policy itself has to search. Trained on 30,000 tool-use trajectories — far fewer than baselines consume — with a regimen aimed at long-horizon tool reasoning, it reports a 20% higher success rate than mainstream baselines in simulation and on real tasks such as pick-and-place in the dark from novel viewpoints, while allowing lighter deployment and incremental addition of new tools.
AgilePE: Autonomous UAV Pursuit-Evasion via Self-Play Reinforcement Learning
Aerial pursuit-evasion demands fast decisions under coupled flight dynamics against an opponent whose behavior keeps changing, which rule-based and differential-game methods handle poorly at high dimensionality. AgilePE trains a policy that maps onboard state observations straight to collective thrust and body rate (CTBR) commands — no trajectory planner or waypoint controller in between — using competitive self-play with Prioritized Fictitious Self-Play (PFSP) against a diversified pool of historical opponents to stabilize optimization. A simulation pipeline modeling actuator response, communication latency, and domain randomization lets the learned policies transfer zero-shot to real quadrotors without task-specific tuning, with hardware runs reproducing the rapid dodging and flanking tactics seen in simulation, including two agents deployed against each other.
Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation
Vision-Language-Action (VLA) models for robot manipulation are almost always benchmarked on static tasks, leaving open how they behave when the world moves and inference latency matters. ReflexBench supplies six dynamic, reaction-critical tasks along with an evaluation harness that decouples simulator stepping from robot control so that configurable latency can be imposed under both synchronous and asynchronous inference. The accompanying ReflexVLA model adds latent future prediction and multi-frame temporal fusion inside the vision backbone and cuts deployment latency with batched visual encoding and CUDA Graph replay, all without large-scale robot-data pretraining. It improves dynamic manipulation while staying competitive on standard static benchmarks, with real-robot experiments under practical deployment conditions.
Expected Free Energy-based Informative Path Planning for Robotic Mars Exploration
A robot surveying unknown terrain — hunting for water on Mars, say — must simultaneously map the environment accurately and reach the highest-value regions fast, all while paying for travel distance and each measurement; classical information-seeking and reward-seeking objectives each optimize only one side. The proposal adopts Expected Free Energy from active inference as a single action-selection criterion, maintaining a Gaussian-process belief over the information field and planning continuous trajectories that minimize expected free energy subject to hard path-length budgets. Across multiple simulated realizations this produces accurate posterior maps and finds top-value regions at the same time, beating information-theoretic baselines under matched settings, with few parameters to tune.
4 more specialized papers
- Adjacency-Based Spectral Proxy Control of Mobile Communication Agents Mariana del Castillo, Federico Larroca
- Learning to Assemble Novel Structures with Unfamiliar Parts under Semantic Constraints Jonghyuk Park, Alex Lascarides, Subramanian Ramamoorthy
- Sensor-Driven Mission Synthesis for UAV/UGV Swarms: A TB-CSPN Coordination Architecture with Hardware-Enforced Safety Uwe M. Borghoff, Paolo Bottoni, Remo Pareschi
- Ensuring Safe Physical AI in Urban Mobility via Hazard-Informed Synthesized Envelopes Alexei Odinokov, Rostislav Yavorskiy