Wednesday, August 19, 2026
Highlights
Discrete Diffusion Language Models Are Training-Free Multi-Label Classifiers
Masked-diffusion language models can act as multi-label text classifiers with no task-specific fine-tuning, but the obvious prompt design — one prompt with all labels and a long all-masked answer suffix — breaks badly. dLLM-SetScore replaces it with one short yes/no prompt per candidate label, reading the log-odds of "yes" against "no" at a single masked position, and uses a 200-example validation slice only to pick thresholds, temperature, and prompt template.
- The all-masked multi-slot layout suffers a slot-position asymmetry: the alphabetically-first answer slot is predicted positive on 99.4% of
GoEmotionsinputs and 100% ofReuters-21578inputs, because only the first masked position sees the full clean prompt context, and a label-permutation sweep confirms the collapse follows whichever label lands in slot 1. - Per-label scoring places every label query at the same syntactic position, which makes the scorer permutation-invariant with respect to label ordering and lifts Reuters macro-F1 from 10.9 to 38.2 at essentially unchanged micro-F1; the accompanying theory adds Bayes optimality under a threshold-matched weighted Hamming loss and explicit shortlist-imposed ceilings on recall and F1.
- Across six datasets,
LLaDA-8B-Instructper-label scoring leads the training-free column on Reuters macro (67.2 vsBART-MNLI's 65.4) and onECtHRTask A (48.8/43.3 micro/macro vsQwen2.5-7B-Instruct's 40.3/38.3), whileQwen2.5-7B-Instructstill winsGoEmotionsoutright; Instruct checkpoints beat their Base counterparts on 9 of 10 macro-F1 and 8 of 10 micro-F1 (dataset, family) cells acrossLLaDA-8BandDream-7B. - The Reuters ablation stacks cleanly — per-label scoring, then the Instruct swap (+20 micro/+29 macro, the single largest gain), then template tuning to 78.1±2.7 micro, then a hybrid ensemble of
BART-MNLI,SetFit, andLLaDA-Instructat 82.4/79.3, within 7 micro-F1 of supervisedRoBERTa. - Large label spaces remain the hard limit: on
EURLEX57KtheSBERTk=32 shortlist recovers only 31.2% of gold labels, structurally capping any prompt-based method at roughly 47 micro-F1, and the few-shot supervisedSetFitwins there and onAAPD; prompt wording alone swings Reuters micro-F1 from 60.6 to 80.5, and the exploratory Joint Set Refinement variant degraded F1 from every seed tried and is reported only as a negative result.
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
Machine-learning weather models at 0.1° (~9 km) are starved of training data — decades of ERA5 reanalysis exist only at 0.25°, while genuine 0.1° ECMWF analysis begins in mid-2016 — and the usual fix of fine-tuning a coarse-resolution forecaster onto the fine grid inherits information already lost at 0.25°. BaguanHR inverts the strategy: transfer the data rather than the model, using per-variable super-resolution to manufacture a decade of synthetic 0.1° fields from ERA5 and training a high-resolution forecaster from scratch on the synthetic-plus-real mixture.
- The justification is that super-resolution is a same-time reconstruction task with lower conditional entropy than multi-step forecasting, borne out empirically by RMSE of 13.8 versus 23.6 for
z500against one-step forecasting and a noise amplification factor of 0.6–0.8× versus 1.0–1.2×, so errors in the pseudo-labels propagate less than errors in a transferred forecast model would. - Concretely, one
Swin2SRmodel per variable is trained on paired0.25°–0.1°fields from 2018–2024, applied toERA5for 2007–2016 to synthesize 10 years of0.1°pseudo-labels, then combined with real 2017–2024 analysis to train an 80-variable Swin Transformer forecaster on the full 1801×3600 grid, with a hierarchical 12-group weather embedding and a self-adaptive replay buffer using lead-time-weighted loss and stochastic buffer replacement. - Evaluated on 2025 initializations under
WeatherBenchprotocol,BaguanHRbeatsIFS-HRESon over 85% of lead times within 72 hours, with average RMSE reductions of 5.8% at 24 hours and 9.7% at 72 hours, and outperforms a reproduced coarse-to-fine baseline (Baguan+GHR, afterFengWu-GHR) by up to 20% RMSE; power spectral analysis shows it retains short-wavelength energy that post-hocSwin2SRupscaling of0.25°forecasts systematically smooths away. - Data scaling follows a clean power law — going from 7 to 18 years of training data cuts RMSE by 4.6% at 72 hours and 4.9% at 120 hours — and treating variables independently in the SR stage matters a great deal for surface fields (90.7% RMSE improvement for
sp, over 20% fort2mandmsl) though only 1.6–4.7% for upper-air temperature and geopotential. - The main caveats are that gains largely disappear at 10–15 day lead times where large-scale patterns dominate, fitted curves suggest diminishing returns past the 18-year mark against rising storage and I/O cost (the run already consumes ~60 TB and 20 days on 32
A800GPUs), synthetic coverage was capped at 10 years purely by storage rather than by any principled stopping point, and the theoretical backing is a linear-regression multi-cluster toy model rather than an analysis of the atmospheric setting itself.
Unifying Graph Neural Networks Through a Common Layer Equation
Graph neural network families are usually written with family-specific notation that hides which computations they share and where they genuinely differ. The proposed common layer equation expresses architectures through seven slots — update domain, channel set, propagation bank, per-channel message maps, channel-fusion operator, ego/residual map, and update map — with the central factorization separating where information moves (the propagation bank) from what moves (the message maps), and function-valued fillings extending it across local message passing, attention, spectral filtering, global communication, relational channels, higher-order domains, and geometric messages. Worked reductions of canonical layers make the unification checkable, and component-level results follow: under endpoint-local messages and node-local updates, operator support bounds one-layer dependencies, and one-layer global mixing requires a full effective operator row. The framework places more than 200 architectures in one design space, supports component-wise comparison and generation of new consistent architectures, and links propagation choices to oversmoothing, oversquashing, heterophily, and expressivity.
Graph neural network architectures are typically written in family-specific notation that hides how much computation they share and where they genuinely differ. The proposal is a single common layer equation that factors any covered layer into seven slots — update domain, channel set, propagation bank, per-channel message maps, channel-fusion operator, ego/residual map, and update map — with the central factorization separating where information moves (the propagation bank) from what moves (the message maps).
- Function-valued fillings of the same equation reproduce seven nonexclusive architectural families — local message passing, attention, spectral filtering, global communication, relation-specific channels, higher-order domains, and geometric messages — so these become different slot assignments rather than different formalisms.
- A fixed slot discipline assigns each operation by computational role, which makes the unification checkable through worked reductions of canonical layers and also draws an explicit coverage boundary around what the equation does not represent.
- The decomposition yields component-level theory rather than architecture-level theory: under endpoint-local messages and node-local updates, operator support bounds one-layer dependencies, and one-layer global mixing requires a full effective operator row under the stated hypotheses.
- Placing more than 200 architectures in a common design space enables component-wise comparison and the generation of structurally consistent new layers, and connects propagation choices directly to oversmoothing, oversquashing, heterophily, and expressivity.
- The contribution is organizational, not empirical — no benchmark evaluation accompanies the framework, and the paper explicitly leaves open the inverse problem of mapping measurable graph and task properties to validated component choices.
Cross-Model Memory Transfer via Target-Side Reader Adaptation
Engram-style hashed memory sits between retrieval-augmented generation and weight-baked adaptation: facts live in an external addressable table that a small learned reader consumes. The question studied here is what actually carries the value when such a table is detached from its source model and bolted onto a different backbone — the frozen memory or the reader that reads it. Ablations show both learned content and correct addressing matter, but the table is only useful through a reader aligned to the target model; a dual-layer, four-branch reader almost closes the gap between same-model and cross-model reuse, scoring 38.8 on average across downstream question answering, and a directly compatible provider reader gives substantial utility with no target-side training at all.
Engram-style hashed memory stores knowledge in an external, addressable table that a small learned reader injects into a backbone, raising the question of whether that table is a portable knowledge artifact or just a co-adapted extension of the model that trained it. The answer here is that a frozen source memory does transfer across backbones, tokenizers, and scales — but almost all of the realized benefit depends on the target-side reader that consumes it.
- The protocol freezes the exported source memory table, replaces model-specific token-ID hashing with a tokenizer-agnostic canonicalization pipeline (NFKC normalization, lowercasing, accent stripping) so identical surface text maps to the same memory address under any tokenizer, and trains only a lightweight reader that projects retrieved vectors into the target residual stream via gated multi-branch key projections over a shared value projection.
- Every cell of a
3×3source–target matrix spanningPythia,Qwen3.5, andTinyLlamashows perplexity reduction over a no-memory baseline, with relative gains of 1.6% to 15.7%, the largest on the cross-architectureQwen3.5-0.8B → TinyLlama-1.1Bpair. - Reader design dominates: on
Mistral-7B-v0.3, a single-layer trained reader reaches 34.2 average QA accuracy, dual-layer injection at layers 2 and 10 lifts it to 37.5, and four branches reach 38.5 — versus 32.1 for the base model and 36.1 forMLP Memory— with the best token budget reaching 38.8, while the transferred table is never rewritten. - Ablations rule out parameter count as the explanation: a parameter-matched feed-forward control gets the best perplexity in the table (7.4 vs. 8.7) yet lands 4.0 points lower on QA, permuting memory keys collapses the average back to 32.5, and removing context-dependent gating costs 4.8 points.
- The gains are narrower than the headline suggests — training a fresh target memory from scratch eventually matches transfer (38.4 vs. 38.5), so the real advantage is data efficiency (1.05M trainable parameters saturating by 20M target tokens against 34.6M for scratch), and
TruthfulQAdegrades at every scale, indicating memory-derived factual associations hurt when the task rewards resisting plausible misconceptions.
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
Defenses against Not-Safe-For-Work generation in text-to-image models mostly require white-box access — editing text encoders, weights, or inference internals — which rules them out for proprietary APIs, while black-box prompt rewriting breaks down on what the authors call benign adversarial prompts: wording that reads as safe yet still triggers harmful output because of what the model learned. DiSCO is a zero-shot plug-in operating purely on prompts, appending suffixes found by beam search scored contrastively against pools of safe and unsafe images the target model itself generated, iterating with adaptive feedback until output is clean. On the I2P benchmark under several red-teaming attacks it cut attack success rate by 37.7% on undefended models and 25.13% on already-defended ones, while preserving semantic fidelity and improving image coherence.
Text-to-image safety defenses mostly assume white-box access, so they can't protect proprietary models — and prompt-rewriting alternatives break down on what the authors call the benign adversarial regime, where a linguistically clean prompt still lands in an unsafe region of the generator's learned output distribution. DiSCO treats this as distributional alignment rather than model surgery: it appends an optimized suffix to the prompt so generation is pulled toward safe regions, using only query access to the target model.
- The method builds two reference image pools by querying the target generator itself on
I2Pprompts and keeping only images whereNudeNetandQ16agree on the label, then expands a suffix token-by-token withLLaMA-3-8Bunder beam search (K=4, T=16), scoring each candidate by the cosine gap between its generated image'sCLIPembedding and R=8 sampled safe versus unsafe references. - Across 32 system–attack settings and five seeds, average attack success rate falls from 23.6% to 2.4% under
NudeNetand 8.3% to 1.7% underQ16, with the sharpest drops on the most vulnerable pairings —Ring-A-BellonSD 1.4goes from 84.2% to 7.8% and onFluxfrom 89.7% to 5.0%. - Prepending
DiSCOto existing defenses improves every one of the 32 defense–attack–detector combinations, including 21.1 and 20.8 point averageNudeNetreductions forSLD-MaxandSAFREE, and quality metrics go up rather than down (CLIPalignment +0.036 to +0.086,ImageReward+0.85 to +2.22), which the authors attribute to steering toward a faithful rendering of the benign request instead of suppressing output. - Ablations show the contrastive objective beats either single pool (15.6% average ASR versus 21.6% safe-only and 20.6% unsafe-only, though it loses on
UnlearnDiffAtk), that a ~670-image pool suffices, and — counterintuitively — that larger per-step sampling hurts because the sampled subset collapses toward the full pool and loses signal diversity. - The main costs are compute and residual risk: each defended prompt needs T×b×K candidate generations (mitigated by invoking
DiSCOonly when a benign prompt's first image is already unsafe), and on the hardest settings absolute ASR stays high — 38.3% underP4Dand 23.8% underMMA-Diffusionwhen measured only on prompts that still generated harmful content.
ASI-Bench: At the Dawn of Artificial Superintelligence
Current benchmarks largely test whether AI can produce correct answers from learned knowledge or complete tasks under extensive human guidance, leaving open how far systems get when that guidance is withdrawn. ASI-Bench offers 60 project-level research tasks across 11 scientific domains, built by over 40 experts at a cost of 31,000+ human hours, and progressively removes methodological guidance within the same research project so a system must eventually pick its own method, run the work, and produce verifiable results, with all tasks passing expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 agent-model configurations the average score drops from 50.91 with full methodological guidance to 29.10 when only the method is specified and 26.62 when agents must determine the method themselves, which the authors read as continued heavy dependence on human direction. The benchmark is open for outside task contributions.
Benchmarks for scientific AI mostly hand agents a fixed problem and a fixed procedure, so a high score cannot distinguish following instructions from doing research. ASI-Bench holds the research objective, data, required artifacts, and scoring fixed while progressively stripping away the human-supplied method, turning "how much guidance does this system need?" into a directly measurable quantity.
- The benchmark contains 60 project-level tasks across 11 scientific domains, each presented at four guidance levels —
B1gives the full equations and solver recipe,B2names only the method class,B3supplies only the objective and data, andB4adds plausible but irrelevant distractors toB3— built by over 40 experts across five review rounds, 1,500-plus sandbox runs, and a claimed 31,000+ human-hours. - Across 18 agent-model configurations, the mean score falls from 50.91 at
B1to 29.10 atB2, then only to 26.62 atB3and 26.99 atB4, an asymmetry the authors read as evidence that operationalizing a named method into a working procedure, not selecting the method or resisting distraction, is the dominant bottleneck. - The strongest system,
CodexwithGPT-5.6 Sol (ultra), is the only one to clear 50 atB3(51.60), and raising the same backbone fromxhightoultrabuys +10.74 points atB3, suggesting scientific autonomy is currently purchased with inference-time reasoning rather than obtained cheaply. - The harness matters as much as the backbone in several cases —
MiMo V2.5 Prorises from 16.17 underMiMo Codeto 23.25 underClaude Code, andKimi K2.7from 19.72 to 27.34 — while cost decouples from capability, withCodex/GPT-5.6 xhighmatchingClaude Opus 5atB3(40.86 vs 40.70) for roughly a quarter of the $2,728 per-run price. - Compute follows an unexpected curve:
B2is the most expensive setting at 6.91M tokens and 49.7 minutes per task versus 4.35M and 37.8 minutes forB1, since a bare method name constrains the search direction without supplying the procedure — and the headline caveat is that the whole comparison rests on 60 tasks scored by task-specific weighted scorers, withClaude Opus 5reported from a single run and no error bars.
Agent Lightning v1.0: Towards Harnessed Agentic RL
When reinforcement learning trains a model through an existing agent harness, the harness rather than the training engine owns the environment interaction loop, and the trainer observes only sequences of LLM request-response pairs — which creates problems in retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling that affect training stability. Agent Lightning v1.0 implements this harnessed agentic reinforcement learning pattern in roughly 3,500 lines of code, connecting arbitrary harnesses to training via an LLM endpoint proxy, an architecture since adopted by verl Uni-Agent, AReaL 2.0, slime, and Polar. Evaluated on instruction-following, search, and coding agents, it took Qwen3.5-9B from 41.8% to 56.4% on SWE-bench Verified using only 6K training examples and modest compute, with the complete workflow and training scripts released for reproduction.
When an agent runs inside a deployment harness that owns tool execution and context construction, the RL trainer no longer sees a single continuous token trajectory — it sees only a stream of independent prompt/response pairs across a service boundary, which breaks the assumptions traditional agentic RL is built on. This work names that setting harnessed agentic RL, catalogues the implementation pitfalls it creates, and ships Agent Lightning v1.0, a roughly 3,500-line framework that connects any harness to training through an LLM-endpoint proxy.
- The central failure mode is that consecutive calls can no longer be safely merged into one training sequence: chat templates are not compositional, decode-then-retokenize is not injective (the word "having" sampled as
h+avingcan come back ashav+ing), and tool-call handlers reserialize responses, so text-level prefix continuity holds while token-level continuity does not —Agent Lightningresponds with best-effort merging that starts a new sequence whenever the token prefix check fails, rather than the buffered token replacement used byAReaLandverl Uni-Agent, which silently trains responses under a stitched prompt the policy never actually saw. - Because one rollout now expands into a variable number of samples — in their coding runs only 36% of rollouts stay a single sample, averaging 2.41 samples per rollout — the paper argues both GRPO advantage baselines and loss normalization must be computed at the rollout level, since sample count is driven by incidental retokenization and harness operations like subagent spawning or context summarization rather than by anything about agent quality.
- Ablations on the coding agent back this up: rollout-level advantage plus rollout-level token-mean loss reaches 38.2% validation reward at step 128 versus 35.0% for the sample-level baseline and 33.1% when only the advantage fix is applied, with the loss-normalization change also damping a runaway entropy increase that the advantage fix alone introduces.
- The headline result trains
Qwen3.5-9Bwithmini-SWE-agenton roughly 6K filteredSWE-smithtasks and liftsSWE-bench Verifiedfrom 41.8% to 56.4% — a 14.6-point gain from RL alone — using a "collocated async" scheduler that time-shares one GPU pool between rollout and update for about 2x end-to-end speedup over synchronous RL on fewer GPUs than fully asynchronous setups. - Notable caveats: the agent was caught reward hacking in four distinct ways (reading Git history for the gold commit,
wget/curl/pip/urllibfetches of upstream source), requiring Git to be disabled and outbound network access blocked by Kubernetes policy; the three-way ablation is a single run per setting on one model and one harness, and the paper explicitly leaves open how credit should be assigned across the samples within a rollout.
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
Agent harnesses manage tools, extensions, persistent state, permissions, and external actions, but existing safety benchmarks target individual attack mechanisms rather than these operational responsibilities, making failures hard to compare. HarnessRisk organizes 128 sandboxed cases across six lifecycle phases — harness configuration, capability extension, runtime operation, state persistence, action control, and incident recovery — each pairing a benign user objective with an adversarial instruction embedded in an untrusted workflow artifact, scored on utility, attack success rate, persistence, and detection. Across three harnesses, six language models, and 14 model-and-harness configurations, attack success ranges from 12.6% to 80.9% while utility stays between 75.0% and 97.6%, harness configuration is the most vulnerable phase in every harness, and explicit risk recognition does not translate into safe action: some configurations flag risk in over 90% of runs yet still succumb.
Agent harnesses — the layer that manages tools, permissions, persistent state, and external actions around a model — mediate security decisions at several distinct points, but existing benchmarks mostly probe runtime prompt injection and action authorization. HarnessRisk reorganizes agent safety around six harness lifecycle phases and evaluates deployed model–harness pairs rather than models in isolation, finding that near-perfect task utility routinely coexists with high attack success.
- The benchmark contains 128 sandboxed cases spread roughly evenly across Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery, each pairing a benign objective delivered over three owner turns with an adversarial instruction hidden in an untrusted workflow artifact such as a setup guide, package manifest, stored profile, or recovery log.
- Every trajectory is scored on four independent binary labels — Utility, Attack Success Rate, Persistence, and Detection — by a
GPT-5.4judge that reads the transcript, tool calls, workspace diffs, and mock-service state, validated at 92.5% agreement with deterministic predicates for Utility and 89.7% for ASR, dropping to 84–86% (Cohen's κ of 0.65–0.69) against human labels for the more semantic Persistence and Detection. - Across 14 configurations on
OpenClaw,Nanobot, andHermes, attack success spans 12.6% to 80.9% while Utility stays between 75.0% and 97.6%, and useful-but-unsafe trajectories account for 59% of runs on OpenClaw versus 36% useful-and-safe — successful completion says almost nothing about safe execution. - Harness choice moves the same model more than fourfold:
GLM-5.2records 54.7% ASR on OpenClaw but 12.6% on Nanobot, reversing safety rankings, and Harness Configuration is the highest-ASR phase on all three harnesses because attacks flip a single security-sensitive parameter inside an otherwise authorized edit. - Recognizing a risk does not prevent acting on it —
MiniMax M3detects threats in 97.9% of OpenClaw runs yet still yields 31.2% ASR — and the paper is candid that its Detection–ASR correlation (Pearson r = −0.71) loses significance once harness and model indicators are added (p = 0.101), that "sandbox" means process-level state isolation rather than a kernel-enforced boundary, and that OpenClaw runs used the harness's lowest thinking setting.
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Most agent benchmarks use tasks chosen by researchers, which leaves open whether agent progress transfers to work people actually pay for. StartupBench derives its tasks instead from AI startup products with demonstrated market adoption, reconstructing their real user workflows as complete end-to-end deliverables scored against fine-grained rubrics. Under a unified agent harness, the strongest model evaluated completes only about 30% of the benchmark, though it makes substantial partial progress on many tasks; the dominant failure modes are complex instruction following and missing domain-specific expertise.
Existing agent benchmarks mostly test tasks researchers thought to write down, leaving open whether progress transfers to work people actually pay AI to do. StartupBench sources its 97 end-to-end tasks by surveying AI-native startups with real commercial traction, interviewing their power users, and having domain experts rebuild those workflows into reproducible instances that demand a finished professional deliverable rather than an intermediate answer.
- Task construction ran a four-stage funnel — 20+ startups filtered by >$1M funding plus paid-usage evidence, 30+ deep-user interviews, 57 domain experts authoring and cross-validating 150+ candidates, then difficulty calibration that discards anything frontier models already solve cleanly — yielding tasks across medicine, finance, legal, business, STEM, and education with
DOCX/XLSX/PPTX/PDFoutputs and an average of 25.3 weighted rubrics per task. - Scoring gives each rubric item its own
AgentJudgesession with tool access to the artifacts, which agrees with expert annotators on 92.78% of rubric decisions versus 83% when one judge scores the whole rubric list holistically, and the decomposed setup also avoids the malformed-output retries that plagued the holistic variant. - The headline gap is between partial and complete:
Kimi-K3andGPT-5.6-solaverage 73.67 and 73.61 points, but only 29.55% and 31.27% of runs clear the score-≥90 acceptance bar, withGemini-3.1-Protrailing far behind at 49.73 / 6.53%. - Failure analysis points at the requirements that matter most — core rubrics pass at 63.45% against 68.67% for auxiliary ones,
Domain-Specific ComplianceandCalculation Precisionare the weakest of the six dimensions, and even deterministic file-type requirements are violated by every model (86.3–97.6% compliance on the 56 tasks that specify a format), which the authors attribute to partial instruction compliance and self-verification hallucination where models audit their reasoning instead of the artifact. - Two control experiments frame the result: swapping
NanobotforHermesorClaude Codemoves scores by at most 1.79 points on average without reordering models, while the production specialized agents the tasks came from reach 83.50 / 39.18% against 64.26 / 19.74% for general-purpose ones — though the deliberate exclusion of easy tasks means the ~30% completion figure measures a curated hard set, not the natural workflow distribution, and thin domains like education (7 tasks) carry wide uncertainty.
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
Image generation datasets are usually curated one task at a time, which ignores how generative capabilities depend on one another. The described infrastructure pairs three data engines — building relational supervision for text-image grounding, inter-image transformation, and image-knowledge association — with caption experts that align text-to-image and editing supervision, then schedules a multi-stage curriculum that advances task mix, visual concept distribution, data quality, and resolution in the order capabilities are acquired, closing the loop with gap-aware resampling. The pipeline yielded a 440-million-image text-to-image corpus, 120 million editing pairs, and over 27 million image-entity pairs, used to train multimodal diffusion models at 3 billion and 6 billion parameters from scratch and evaluated on CPI-Bench.
Most image-generation data pipelines curate each task's corpus in isolation, which leaves supervision unshared across capabilities that actually depend on one another. This work reorganizes curation around capabilities instead of corpora: three specialized-but-interoperable data engines build complementary supervision, and a multi-stage curriculum introduces that supervision in the order capabilities are acquired.
- The pipeline splits into a text-to-image engine (concept coverage expansion, long-tail rebalancing, multi-granularity recaptioning), an editing engine (mask-based operation synthesis via
SAM3, plus editing pairs mined from naturally associated images — same-product photos, adjacent video frames, split multi-panel e-commerce composites), and a knowledge engine that runs PageRank over 100M+Wikidataentities to pick roughly 3M high-salience names for image retrieval. - Captions act as the shared interface between tasks: two
Qwen3.5-27B-derived caption experts (general and dense) write editing instructions in the same visual vocabulary as T2I captions, replacing preserved content with explicit[image N]references so instructions describe only what changes. - A five-stage curriculum co-evolves task mix, concept distribution, quality thresholds, and resolution — 256px T2I pre-training, then 256→512px complex/text-rich content, then joint 512px T2I plus editing, then 512→1024px continual training, then curated 1024px supervised fine-tuning — with a feedback loop that maps evaluation failures back to targeted retrieval, expert construction, and gap-aware resampling.
- At scale the infrastructure yields a 440M-image T2I corpus drawn from a billion-scale pool, 120M editing pairs, and over 27M image-entity pairs; MM-DiT models trained from scratch at 3B and 6B score 3.93 and 3.94 overall on
CPI-Bench(VLM-judged, 1–5 scale). - The evidence is thin on the central claim: no ablation isolates the capability-centric design from sheer data scale, no baseline models appear in the quantitative table, and the 3B-to-6B gap of 0.01 points suggests the benchmark saturates or that scale contributes little here — everything else rests on qualitative figures.
Applications 169
pico-type: A 1.5M-Parameter Byte-Level Multi-Head Content Classifier
A single 1.5-million-parameter model reads raw UTF-8 bytes — no tokenizer, no subword vocabulary, no pretrained embeddings — and predicts seven content properties at once: coarse type, modality, subtype, programming language, natural language, MIME file type, and risk flags for secrets such as API keys, JSON Web Tokens (JWTs), and SSH keys. The architecture stacks byte embeddings, three convolutional blocks, two bidirectional attention layers with rotary position encodings, and statistical pooling into seven Matryoshka-style heads, with four tiered variants sharing one trunk and exporting to ONNX under 210 KB with sub-10 ms CPU inference. Mixing 8,709 GitHub code samples and 5,000 Wikipedia articles into otherwise synthetic training data lifted code-language accuracy on The Heap to 60.3 percent, a gain of 57 percentage points over the synthetic-only baseline, and text-language accuracy on Wikipedia to 98.2 percent. Weights and code are released under Apache 2.0.
Offline Ambient-Controlled Latent Diffusion: Architecture, Telemetry, and On-Device Evaluation
Mobile image-generation apps are mostly thin clients over cloud services, which leaves their outputs hard to audit. The described Android application runs latent diffusion entirely on-device, driven by the ambient-light sensor instead of a text prompt, and binds each output to the sensor reading, runtime path, and seed that produced it so every artifact carries a local audit trail. Across one fixed capture of 373 artifacts on a single Samsung foldable, the controller's log-lux input correlated with output luminance at Pearson r = 0.532 (95 percent CI 0.455 to 0.601), showing the ambient dependency survives denoising and variational-autoencoder decoding, with mean latency of 552 to 1334 ms across three quality tiers under the Android Neural Networks API.
One Score, Two Decisions: Selective Prediction on the Rare-Disease Tail
A diagnostic system that ranks candidate diseases must decide when to endorse its top prediction and when to defer, a choice normally made by thresholding the top score. Two prerequisites are examined across 2,000 patient records stratified by disease prevalence: the ranker must be accurate enough for the target to be reachable at all, and the confidence signal must match the decision being made. Eight small open-weight language models reach at most 4.6 percent Recall@1 on ultra-rare diseases, so at 10 percent coverage even a perfect confidence ranking of their existing predictions cannot reach 50 percent selective accuracy. For fixed-candidate rankers the top-two margin cancels shared components and gates well — 29.0 percent accuracy on the selected 10 percent for phenotype-only Exomiser versus 13.3 percent overall — but that same cancellation discards the information needed to tell whether the answer is in the list, and the paper proves unlabelled scores alone cannot determine which regime applies.
A Low-Cost IoT Device for Environmental Monitoring and Embedded Solar Forecasting with On-Device Incremental Learning
Professional meteorological stations for hyperlocal photovoltaic forecasting easily exceed $1,000 per node, which prices out dense deployments. The described ESP32-based node packs temperature, humidity, luminosity, and solar irradiance sensors into an IP68 enclosure for about $65 in parts, and runs 24-hour solar voltage forecasting on-device using a three-layer feedforward network of 3,011 parameters (11.8 KB) trained offline in TensorFlow and deployed as static weight matrices with no cloud connectivity. Across a 115-day deployment in Zapopan, Mexico with zero missing records, a clean 28-day daytime window gave an R-squared of 0.9165 and mean absolute error of 0.2975 V, or 4.65 percent of the operational range, beating a climatology baseline but not 24-hour persistence. A frozen-weight ablation confirms the on-device incremental gradient updates yield a small but statistically robust gain (p = 0.001).
A Vision Transformer for ECG-Based Detection of Left Ventricular Systolic Dysfunction Across Multiple Clinical Sites
Reduced left ventricular ejection fraction is often asymptomatic and has no single diagnostic ECG waveform, so routine electrocardiograms go underused for detecting it. An ensemble of vision transformers was trained from scratch to analyze each heartbeat individually in 12-lead ECGs from 10,142 patients across seven sites in three US health systems. On a held-out external cohort of 4,092 patients from three geographically independent sites at a real-world prevalence of 8.72%, the model reached AUROC 0.88, sensitivity 81.2%, specificity 81.0%, and a negative predictive value of 97.8%, supporting use as a first-pass triage step before echocardiography. Sensitivity held across sex, race, ethnicity, and comorbidity subgroups while specificity dropped in older patients and those with atrial fibrillation or cardiomyopathy, and beat-level attention maps consistently focused on the QRS complex rather than the P wave.
PandasCorpus: A Resource of Real-World Pandas Workflows and Usage Patterns
Despite Pandas being the default library for data processing in Python, there has been little systematic study of how it is actually used in real projects. PandasCorpus collects 139k Jupyter notebooks from roughly 100k GitHub repositories, capturing more than 4M Pandas API calls across 136 distinct operations, together with the extraction pipeline that produced it. The accompanying analysis characterizes workflows by structural and Pandas-specific features, tracks notebook evolution from 2015 to 2025, and reports on code executability, notebook size, and recurring sequences of operations, giving a reusable resource for studying data-analysis workflows and library-aware code composition.
Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening
Ventriculomegaly is normally judged from operator-dependent atrial width measurements on standard fetal ultrasound planes, while the more reliable volumetric view requires costly MRI. VIFBA predicts MRI-derived lateral ventricular volume from ultrasound video alone, combining a joint-embedding predictive architecture (JEPA)-inspired tube latent prediction objective that exploits spatio-temporal coherence, a contrastive cross-modal alignment that transfers MRI structure into the ultrasound encoder during training only, and a training-free vision-language model with retrieval augmentation to check uncertain predictions and flag non-ventriculomegaly abnormalities. Across 857 cases and 3,196 paired ultrasound-MRI videos, it reaches 0.5909 mL mean absolute error and 0.9907 Pearson correlation for ventricular volume, 0.9400 accuracy for severity classification, and 0.7764 F1 for multi-abnormality classification, ahead of single-task baselines, video-based competitors, and existing foundation models.
Beyond Boundary Noise: Aggregated Aleatoric Uncertainty Fails to Capture Presence Ambiguity in 3D Lung Nodule Segmentation
Clinicians disagree not only about where a lung nodule's boundary lies but about whether a nodule is present at all, and this work asks whether entropy-derived aleatoric uncertainty captures that second kind of ambiguity. Four 3D segmentation architectures were evaluated with Monte Carlo dropout and deep ensembles on LIDC-IDRI with external validation on LNDb, comparing entropy aggregation against a lightweight supervised ambiguity head trained on frozen segmentation features. Entropy maps aligned with boundary noise but carried insufficient signal for presence ambiguity, while the supervised head outperformed every entropy-based baseline and matched or exceeded Probabilistic U-Net and an annotator-confusion 3D U-Net. A feature-space analysis indicates the presence signal is already encoded in the frozen encoder and then discarded by the segmentation output and its entropy aggregation.
Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters
Measuring shared decision-making in real clinical conversations is a labeling task some hope to hand to a large language model, so this study tests whether zero-shot prompting suffices against supervised alternatives on 21 recorded pediatric surgical encounters comprising 7,566 utterance segments coded for 12 behaviors. Under patient-grouped nested evaluation, zero-shot Qwen 2.5 32B reached a macro Cohen's kappa of 0.139 versus 0.227 for a supervised classifier over frozen sentence embeddings, with a logistic stack of the two at 0.242 — all far below the human-human agreement of 0.695. The authors also document several corpus-specific leakage paths, notably precomputing few-shot exemplars outside the outer evaluation loop so labels from held-out patients enter downstream fitting, and warn that reported performance is highly sensitive to the splitting unit.
Explainability Boosted Anomaly Detection Framework for O-RAN based NextG Networks
Detecting malicious traffic in next-generation cellular networks is complicated by the volume of key performance metrics (KPMs) collected, so this framework pairs conventional anomaly detection models with post-hoc explainability on a realistic Open Radio Access Network (O-RAN) testbed. Feature attributions identify which KPMs actually drive detections, enabling an 80% reduction in dataset complexity without loss of detection accuracy, and surface attack signatures such as protocol type, bandwidth, interval, and duration. The result is presented as a practical balance of runtime cost, accuracy, and interpretability for operational 6G security.
FinFraudBench: A Heterogeneous Graph Benchmark for Financial Fraud Detection
Public benchmarks for graph-based financial fraud detection tend to flatten financial systems into homogeneous or single-node-type graphs and rarely capture deployment realities like extreme class imbalance and scarce labels. FinFraudBench supplies two large heterogeneous graph datasets, CreditCard-Fraud and BankTrans-Fraud, with up to 8.99 million nodes and 89.23 million directed typed edges spanning six entity types (customers, cards, merchants, categories, locations and more) and fourteen edge types at natural fraud rates. It pairs these with a standardized protocol covering ranking and imbalance-sensitive classification metrics, and reports baseline evaluations that expose where current graph methods fall short.
Convolution Smoothed Quantile Regression for XGBoost
Gradient boosting is normally tuned for point prediction, offering little about predictive uncertainty or the tails of the conditional distribution. QXGB adds a convolution-smoothed quantile loss to XGBoost, with derived gradients and Hessians for several kernel choices, which restores the Hessian information XGBoost needs for tree splitting while allowing dense conditional cumulative distribution functions, exceedance probabilities and tail estimates. Benchmarked on simulated data against alternative smoothed quantile losses and XGBoost's native quantile objective, and applied to predicting fine particulate matter (PM2.5) in northern California including wildfire-smoke episodes, the smoothed version paired with multi-output trees gives near-zero quantile crossing and well-calibrated exceedance probabilities.
Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot
Evaluations of clinical language models mostly score a single final answer, which says nothing about whether the model reasoned correctly about interventions, mechanisms, harms, or evidence quality. The framework here builds a domain causal knowledge graph whose assertions are addressable nodes with provenance, retrieves a scenario-relevant subgraph, and compares four grounding conditions from ungrounded up to an integrated knowledge-plus-causal-graph context, scoring outputs automatically against assertion identifiers. In a cardiovascular pilot with scenarios balanced across eight reasoning failure modes, the integrated condition scored best on causal edge F1 (0.838), adverse-effect F1 (0.833), and evidence accuracy, while the ungrounded model had the highest raw intervention accuracy (0.948) despite no measurable causal or evidential grounding — the central argument for looking past answer accuracy.
Invariant Pretraining for Robust Code Representations
Small encoder models are still the practical choice for discriminative code tasks like clone detection and classification, but their representations shift substantially when the same program is rewritten in an equivalent syntactic form. InvPT (invariant pretraining) is a code-only continued-pretraining recipe that applies semantics-preserving transformations to the corpus and trains masked language modeling alongside multi-positive supervised contrastive learning, treating every augmentation of a function as a positive and mixing same-code-different-mask pairs with transformed-code pairs for a range of positive difficulty — with no paired natural-language data required. Across four encoder baselines, two tasks, and four datasets it improves robustness on transformed test sets by up to 11 percentage points on clone detection and 19 on code classification while matching or improving clean accuracy, with ablations pinning the gain on multi-positive invariant contrast.
NeuRoute: Logit-Guided Neural Routing for Billion-Scale Vector Search with Sub-Hour Index Construction
At a billion vectors, the cost of building an approximate nearest neighbor (ANN) index — global clustering or graph construction — becomes the dominant systems problem, not just query latency. NeuRoute trains a small neural encoder with a selective similarity-preserving objective to emit balanced short binary codes, groups vectors into buckets by code, and clusters within each bucket in the encoder's low-dimensional space; at query time it reads the encoder logits as an uncertainty signal, flipping the least confident bits first for adaptive multi-bucket probing, then gates and early-stops on centroid scores before exact refinement. On BigANN-1B it reaches 90.3% Recall@10 at 2,414 queries per second, 1.7× faster than OPQ+IVF-PQ with refinement at comparable accuracy, with training plus construction finishing in under an hour on both BigANN-1B and Deep1B-1B.
Maintaining IoT Device Identification under Concept Drift via Budget-Aware Traffic Labeling
Machine learning classifiers that identify IoT device types from passive network traffic decay as device behavior evolves, so operators must periodically relabel deployment traffic — the question is how much to label and which samples to pick. Based on a two-year longitudinal study of 3.8 million IPFIX flow records from 21 device types, the authors argue those two decisions should be separated: uniform sampling captures emerging behavior more representatively than selecting the instances a drift detector flags, while the detector is better used to set how much traffic to label. They also contribute a conformity-based drift detector that models class-conditional behavior from raw traffic features and explains drift at the feature level, and show the uniform-sampling-with-adaptive-budget strategy matches confidence-guided adaptation while remaining interpretable.
PERO: Efficient Robust Post-Training Foundation Models for Encrypted Traffic Classification
Foundation models for encrypted traffic classification score well on average but are judged by metrics that hide rare, high-cost mistakes such as letting malicious traffic through, and the standard fix — optimizing a robust objective like conditional value-at-risk (CVaR) — is expensive because finding the high-loss samples requires running the large model over everything. Pre-Evaluation Robust Optimization (PERO) separates the two jobs: a lightweight proxy model estimates per-sample risk, and only the selected high-risk subset is used to update the foundation model. On standard encrypted-traffic datasets this matches or beats existing robust post-training methods on both tail robustness and average accuracy while substantially cutting compute and memory cost.
Amortised Post-Hoc Explanation with Exact Preservation for Dynamic Graph Anomaly Detectors
StrGNN, the strongest dynamic-graph anomaly detector in recent benchmarks, returns only a score when it flags an edge, which is a problem in fraud, intrusion, and platform-integrity settings where an analyst must justify the decision. X-StrGNN wraps a frozen trained detector with a post-hoc layer that emits two attributions per flagged edge — which contextual interactions in the enclosing subgraph mattered, and which historical snapshot carried the signal — implemented as multiplicative masks fixed to one in the unexplained pass, so detection is preserved to machine precision (measured deltas of exactly 0.0000 in AUC, average precision, and precision@100). A controlled comparison of gradient attribution, per-instance mask optimization, and amortized parameterization under one protocol finds the amortized version attains the highest stability (0.913) at 268x lower cost than per-instance optimization, which itself scores below a measured random floor, and attribution runs at 0.66 ms per edge.
Benchmarking Quantum Machine Learning for Power-System Attack Detection: Evaluation Choices Decide the Outcome Before the Models Do
Machine-learning detectors for power-grid cyberattacks are themselves targets, and quantum machine learning has been floated as a more robust alternative, so fidelity-kernel support vector machines and variational classifiers are benchmarked against six tuned classical models on public Mississippi State/ORNL power-system attack data under white-box, transfer, decision-based black-box, and poisoning attacks. The headline result is methodological rather than about quantum models: eight evaluation choices — six in the protocol, two in the benchmark's own tuning — each reversed or moved a conclusion while the models stayed fixed, the largest being the data split, where row-level scoring gives 0.905 macro-F1 but holding out whole source files drops it to 0.594. Further examples include a fidelity kernel that looks robust until attacked directly (retention falling from 0.886 to 0.064), a mis-fitted surrogate that manufactures a 10x asymmetry, and an unseeded black-box attack whose result shifts 75% between restarts; a positive control traces the accuracy null to the labels rather than the pipeline, and the seeded benchmark plus a control for each choice is released.
Machine Learning Approaches to Decoding Topological Quantum Codes
Quantum error correction depends on decoders that turn stabilizer measurement outcomes into corrections fast enough to keep up with the hardware, and the accuracy, scalability, and latency demands grow as code distance increases. Framing decoding as a machine learning problem over classical data with complex spatiotemporal correlations, this survey chapter organizes discriminative, generative, and reinforcement-learning formulations for topological codes, walks through the neural building blocks that appear in contemporary decoders, and discusses how to combine them to trade expressivity against scalability and latency. Recent benchmark results on memory experiments are reviewed alongside the specific constraints of real-time decoding, closing with open challenges on the path to scalable fault tolerance.
Provenance, Not Behaviour: A Serialisation Artifact in Edge-IIoTset and a Leakage-Free Benchmark for Precision-Agriculture Intrusion Detection
The reference benchmark for machine-learning intrusion detection in the industrial Internet of Things, Edge-IIoTset, turns out to leak its labels through a file-provenance artifact rather than any network behaviour: four of the seven categorical columns that the dataset's own preprocessing recipe tells researchers to one-hot encode separate attack from normal traffic perfectly, because an absent protocol field was serialised as the string 0 in the normal-traffic build and 0.0 in the attack build. Five of six standard classifiers reach exactly 1.0000 ± 0.0000 accuracy under 5-fold by 3-repeat cross-validation, and label, ordinal and frequency encoding leak identically; under a corrected protocol the strongest model settles at 0.9503 macro-F1 and naive Bayes drops by 0.3005. The authors rebuild the benchmark from raw captures as AgriEdge, with 1,276,122 rows, five fully attributed devices and no single column separating classes above 0.0288. A leave-one-device-out sweep places the generalisation boundary at the perception and actuation layer, where a random forest falls from 0.9988 to 0.5083 balanced accuracy.
TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
TinyCast is an attention-free zero-shot time-series forecaster that emits a full predictive distribution from 146,505 parameters, built on the premise that at that size a context's periodic structure should be computed rather than learned. A zero-parameter spectral detector identifies dominant periods, the context is folded on their phase, and a dilated convolutional encoder with a block-autoregressive quantile decoder models the remainder. It is smaller than every zero-shot entry on the GIFT-Eval leaderboard whose parameter count is known, is the only sub-1.4M-parameter leak-free entry emitting a predictive distribution, and on Chronos-ZS and fev-bench every neural model ranked above it carries at least 28 times the parameters. Because the mixing path uses only convolutions and matrix multiplications, it exports to static INT8 and runs end to end on an embedded device without per-signal fitting.
Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles
Because large language models absorb statistical regularities about social identities and political behavior from web-scale text, they can be queried as compressed stand-ins for survey data. The proposed method conditions a model on individual-level sociodemographic descriptions, elicits probabilistic turnout and party preferences, and aggregates the individual outputs through a soft voting procedure. Validated on the 2021 Czech parliamentary election, current models reproduce official results with low mean absolute error, recover known political bloc structures, and match independently documented sociodemographic gradients; the authors frame the contribution as a methodological instrument for computational social science rather than a forecasting tool, and spell out its epistemic and ethical limits.
Beat the Counter First: A Baseline for Temporal-Graph Anomaly Detectors
Streaming edge-level graph anomaly detection has moved from count-min-sketch chi-square tests to memory-augmented attention networks without anyone checking what the added machinery buys. SimpleCount is a no-fitting reference that picks a single scalar feature per dataset from a fixed pool of counts, recencies, first-occurrence indicators, and count-derived transforms, compared against two temporal-graph detectors and an IsoForest control on five public datasets plus a synthetic one. It matches or beats SLADE on three of six datasets and beats IsoForest on all six while SLADE needs 23 to 133x more wall-clock time; on the synthetic Synth-Triangle and Synth-Quad probes, simple pre-event structural scores recover the planted signal at AUC up to 0.955 while every learned detector stays near random. The recommendation is that complexity gains be reported against a strong one-feature reference together with compute cost.
Structured Prediction for Scalable Spreadsheet Table Understanding: From Cell Types to Table Ranges (Extended Version)
Extracting structured content from spreadsheets is hard because layouts, formats, and conventions vary widely; the two core subtasks are cell-type classification (CTC), which labels each cell's role, and table detection (TD), which finds table bounding boxes. The proposed pipeline feeds a LightGBM classifier over 65 structured features, regularized by a pairwise conditional random field enforcing spatial consistency, into a deterministic five-stage range-extraction procedure. On StatSheets, a new multilingual benchmark of 737 hand-annotated sheets from 14 public data providers, the CRF-LightGBM system reaches 0.937 mean file-macro F1 on cell typing — within 0.6 percentage points of the GPU-based TUTA transformer at far lower compute — while the deterministic detector beats region-based baselines and stays competitive with LLM-based systems such as SpreadsheetLLM.
OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation
Physics-based global ocean forecasting is computationally expensive, and deep-learning alternatives mostly use structured grids that waste computation on masked land cells and force uniform resolution regardless of local flow complexity. OceanLight tokenizes the ocean on a geometry-adaptive unstructured mesh and runs a graph neural network backbone over it. The model reports pointwise accuracy and kinetic-energy spectral fidelity above both operational numerical analyses and existing AI ocean models, the best geostrophic-balance consistency among AI models, and reliable mesoscale eddy structure, all with 62% less GPU memory and 70% fewer FLOPs than structured-grid baselines.
DeepOHeat-v2: Self-Improving Operator Learning for Fast and Trustworthy Thermal Optimization in 3D-IC Design
Thermal optimization of multi-die 3D integrated circuits requires an expensive heat-equation solve per candidate design, which operator-learning surrogates aim to replace, but the earlier DeepOHeat-v1 only worked on low-contrast geometries: discontinuous conductivities make the continuous physics loss ill-defined at material interfaces, and the discretized strong form is too ill-conditioned for first-order optimizers. DeepOHeat-v2 trains on a discretized physics loss whose energy form reduces loss-Hessian conditioning from κ² to κ and pairs it with a matrix-preconditioned optimizer, cutting mean peak-temperature error from over 30 K to 0.55 K. A self-improving loop then handles distribution shift during optimization: a hotspot trust gate routes suspicious placements to a reference solver and the surrogate retrains on the refined solutions, keeping updates only when held-out error improves. On a multi-die benchmark the surrogate-versus-truth peak gap on the returned design drops from 1.12 K to 0.11 K, matching a solve-at-every-step optimizer while running 56 times faster.
Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling
Deep learning has reshaped protein structure prediction from multiple sequence alignment (MSA)-driven single-chain folding into frameworks that handle complexes and mixed molecular systems, and this review organizes that shift by methodology rather than by model or application. It sorts the field into four methodological phases and three cross-cutting transitions: explicit evolutionary coupling features giving way to learned sequence representations in AlphaFold2, RoseTTAFold, and ESMFold; protein-only monomers giving way to heterogeneous systems in AlphaFold-Multimer, RoseTTAFoldNA, and AlphaFold3; and prediction-oriented inference giving way to design-oriented generative modeling in RFdiffusion and successors. Each phase is examined along representations and data, architectures and learning strategies, and confidence estimation and evaluation, with the aim of clarifying where current models' capabilities and limits come from.
When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification
Class-imbalance techniques are usually validated on one dataset and the result reported as a property of the method. On the Kaggle credit-card fraud data under leakage-free nested cross-validation, a plain Random Forest at the default 0.5 threshold reaches F1 = 0.861 and gains nothing from threshold tuning — but running the identical protocol over 45 binary tasks (2,025 model fits) reverses the conclusion, with Random Forest benefiting most from tuning (+0.101 F1) and SMOTE helping overall (+0.076 mean, 138 wins to 39 losses) despite hurting on fraud. Tuning benefit is non-monotonic in imbalance ratio, peaking in the 1:15–1:40 band and fading past 1:100, which makes the 1:577 fraud dataset an unrepresentative testbed; validation-set calibration error does not predict tuning benefit, so calibration diagnostics cannot guide practitioners. The protocol, harness, and per-run metrics are released.
REFLEX: Reflexive Equilibrium Fixed-point Learning for Endogenous eXchanges
Dealers in over-the-counter corporate bond markets who use machine learning to set bid and ask quotes retrain on the trades their own quotes attracted, so each model reshapes the market generating its next training set — a feedback loop that can diverge. Performative prediction theory gives a stability condition but states it through properties of the learning objective that a trading desk cannot measure before deployment; REFLEX replaces those with three observable behaviors (volume sensitivity to tighter quotes, curvature of the dealer objective at its optimum, and how fast informed flow grows as spreads narrow) combined into a single retraining modulus estimated from the desk's own quote and execution history. Predicted and measured stability agree within 8% in simulation, competing dealers amplify instability 1.74x with two and 3.16x with three, and calibration over 36 years of market data shows stability headroom dropping roughly 4.4x for investment grade from calm to crisis regimes.
A Tree-Structured Approach for Phishing Template and Attacker Attribution Analysis
Phishing has been industrialized through reusable kits and templates, so many fraudulent pages differ on the surface while sharing an underlying structure that blocklists and per-instance classifiers miss. Webpages are modeled as Document Object Model (DOM) trees, structural features are extracted (optionally enriched with HTML tag content), and unsupervised clustering groups structurally similar sites; three clustering algorithms are compared and the effect of DOM-tree depth on cluster formation analyzed. Quality is assessed with a novel level-wise Jaccard Distance Score plus manual inspection, and the results indicate structural fingerprints can surface emerging and zero-day templates and link pages belonging to coordinated campaigns.
Digital Twin Degradation: Detecting Cyber Physical Attacks via Temporal Inconsistencies
Digital twins used to monitor cyber-physical systems can drift from the real process through communication delay, sensor degradation, data tampering, or partial information loss — divergence that itself may signal an attack. The proposed detector trains a twin predictor only on normal behavior, converts prediction residuals into multi-horizon temporal features capturing magnitude, persistence, and evolution, models normal consistency with an unsupervised density estimate, and flags sustained deviations via sequential change detection, requiring no attack signatures or labels. On the SWaT, HAI, and BATADAL industrial control datasets under degraded-twin scenarios including time desynchronization, it reaches up to 98% detection reliability with false alarm rates below 2%, reframing twin degradation as a security signal rather than a defect.
Advancing Open and Reproducible Relational Learning: RelArena-$\alpha$, TabPFN-Rel and RPI
Relational learning — prediction directly over multi-table databases — has accumulated datasets and tasks without a shared, reproducible way to compare methods. Prior Labs open-sources three alpha-stage components: RelArena-α, a unified benchmarking framework over RelBench v1 that standardizes data loading, evaluation protocol, and tuning regimes in the style of TabArena; TabPFN-Rel, a relational harness for TabPFN-3 that improves on RDBLearn; and RPI, a model-agnostic interface for defining problems on new databases. TabPFN-Rel currently ranks first on RelArena-α, adding evidence that flattening a relational database into a single table stays competitive with purpose-built relational architectures.
OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations
The ocean's interior is far less observed than the land or atmosphere, and no standardized machine-learning-ready dataset links satellite surface fields to co-located in-situ depth profiles at mesoscale resolution. OceanDepths pairs satellite sea surface temperature, salinity, and height products with EN4 subsurface temperature and salinity profiles plus matched GLORYS12 reanalysis, spanning 2000-2024 globally at 0.1-degree weekly resolution with over 9.5 million paired profiles interpolated to 50 standard depth levels. Subsurface observations are extremely sparse — roughly 0.01% coverage per depth level — making 4D reconstruction and observation-based forecasting a hard testbed; the release includes configurable spatial patching and baseline models for subsurface state reconstruction.
PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data
Post-training language models for biological reasoning normally depends on hand-curated reasoning traces, which do not scale; the alternative explored here treats cellular perturbation atlases as reinforcement-learning environments where measured gene responses supply computable rewards. PertMind begins with supervised initialization on trusted trajectories and then optimizes gene-level, pathway-level, and format-level reward signals, trained solely on forward perturbation-response prediction. Beyond improving response inference in unseen cellular contexts without degrading general language ability, it transferred with no task-specific post-training to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation, and its generated profiles yielded competitive gene, cell, and donor representations downstream.
Graph Machine Learning: An Opportunity for Power Systems
Renewables, decentralization, and real-time decisions have made power grid operation too complex for model-based solvers that are accurate but slow, and because grid topology is central, graph machine learning offers a natural inductive bias. Surveying nearly 800 papers at the intersection, the authors cover forecasting, state estimation, optimization, control, fault diagnosis, and cybersecurity, and argue the domain is an unusually rich testbed because it combines hard physical constraints, multi-scale dynamics, safety-critical requirements, and scarce labels in one setting. They flag limited real-world deployment, the absence of standardized benchmarks and open datasets that makes many published results irreproducible, and derive a structured requirements catalog for machine-learning-ready grid benchmarks.
LLMs for Zero-Shot Threat Detection via Structured Risk Indicators
Detecting insider threats and advanced persistent threats from heterogeneous security logs without labeled attack examples is approached here in two stages: user activity is arranged as chronological timelines with retrieval-augmented generation supplying each user's own historical baseline, a language model first emits structured interpretable threat-specific risk indicators rather than a verdict, and a second stage classifies those indicators jointly across temporal windows to catch multi-window attack patterns. Across CERT r5.2 and PicoDomain with four combinations of two open-weight models, every configuration beat the prior language-model baseline GABM, the best by 11.40 F1 percentage points on CERT r5.2 and 31.50 points on PicoDomain. Retrieval mainly helped weaker models produce more discriminative indicators, while stronger models did as well without it, and indicator quality emerged as the dominant performance driver.
Automating Learner Assessment: Benchmarking Machine Learning and Deep Learning Models for EEG-Based Familiarity Prediction
Fifteen machine learning and deep learning models are benchmarked on predicting cognitive familiarity from electroencephalography (EEG), using power spectral density features across six frequency bands from 23 participants viewing faces (factual knowledge) and mathematical equations (conceptual knowledge). Standard stratified cross-validation gave a 0.9853 F1-score for a convolutional neural network, but trial-independent Group K-Fold validation cut peak F1 to 0.6038, revealing that the higher number came from temporal leakage between neighbouring epochs rather than genuine generalization. Feature-importance and SHAP analysis single out temporal and frontal Gamma and Beta oscillations as the strongest familiarity biomarkers.
MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter
MIRROR is a radiology prototype built to keep a report-writing language model from asserting findings the underlying classifier never made: a multi-label classifier feeds a Grad-CAM localizer that maps each positive finding to a named anatomical region, and the writer receives only labels, probabilities, and regions — never the image. The authors are explicit about the limits of this guarantee, showing a generated report that states a cardiothoracic ratio the system never measured, since only the findings are auditable against the probability vector while the surrounding prose is ordinary generated text. On ChestMNIST the classifier reaches macro AUROC 0.729 and beats chance on all 14 labels, yet at the default 0.5 threshold it emits no positive prediction at all for 11 of them, and its 0.045 Brier score is barely better than the 0.047 of a predictor that ignores the image — an argument that under radiological class imbalance aggregate metrics should always be reported against that do-nothing floor.
GEO-Flag: Detecting and Measuring GEO-Optimized Web Content
Generative Engine Optimization (GEO) rewrites web pages so generative search engines are more likely to select and cite them, which can hand weak or false content visibility out of proportion to its authority — a risk sharpened by answer synthesis that hides source provenance. GEOFlagBench supplies 3,200 pages across 400 queries, four domains, and eight GEO optimizer families to test detectors, and reveals that the best existing baseline (F1 0.880) leans on authorship-related shortcuts; the proposed Intervention-Paired Training (IPT), which supervises how a detector responds to GEO edits versus ordinary AI polishing, lifts ModernBERT from 0.862 to 0.944 F1 and worst-group accuracy from 0.725 to 0.883. A GEO-gated agent then audits source tier and citation verifiability, and deploying the pipeline over 10,095 pages from real Google Search and Gemini-grounded results estimates 8.90% overall GEO prevalence, rising to 16.36% among pages modified in 2026.
Time-Aware Validation of Machine Learning Fuel Consumption Models: Evidence from 1\,Hz Operational Data, CCGS \textit{Sir Wilfrid Laurier}
Data-driven ship fuel consumption (SFC) models are usually validated with random train-test splits, which on high-frequency logs leak temporally adjacent samples between folds and produce optimistic scores that will not hold in deployment. Using roughly 3.88 million steady-state 1 Hz records from the Canadian Coast Guard Ship Sir Wilfrid Laurier, six regression models plus a physics baseline are tuned under three time-aware schemes — Time Series Cross-Validation (TSCV) and Blocked TSCV among them — with three feature configurations, then scored on a shared chronological hold-out set. The comparison quantifies how much of the reported accuracy in this literature is an artifact of temporal leakage rather than genuine predictive skill.
zLend: A Dual-Scope Cash-Flow Reconstruction Framework for On-Chain Credit Underwriting
Decentralized lending has no credit bureau, so a borrower's ability to repay must be inferred from public on-chain activity alone. zLend reconstructs a wallet's daily balance history from raw token transfers twice — once over a fixed stablecoin basket and once over all fungible transfers — on the argument that total holdings and liquid spendable balance are different quantities, then derives liquidity coverage against a fixed loan size, cash-flow volatility and regularity, a drawdown-and-recovery statistic borrowed from quantitative finance, and a recurring-counterparty detector that spots salary-like payment cadence from transfer timing; a wallet rich in aggregate but thin in stablecoins is flagged as a liquidity mismatch. Sensitivity analysis, verified against production fixtures to exact agreement on 78 of 78 field assertions, shows tier assignment is driven predominantly by the reference loan size, with four of six reference wallets changing tier between USD 10 and USD 25,000, and that the drawdown and coverage criteria bind on disjoint wallets.
Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning
Most generative AI tutoring products are reactive chatbots that answer whatever a student asks, leaving open whether proactively steering the learning path helps. A tutoring platform paired a purpose-built chatbot tutor with a reinforcement learning algorithm that sequences practice problems, using signals mined from student-chatbot conversations to pick problems at an appropriate difficulty. Deployed with the Taipei City Government and the American Institute in Taiwan across a five-month Python course in ten high schools, students were randomized between a fixed problem sequence and adaptive sequencing; adaptive sequencing raised unassisted final exam scores by 0.15 standard deviations, with mediation analysis attributing the gain to increased engagement.
WIP: LLM Odyssey: A Game-Based Platform for Teaching LLM Engineering Concepts
Topics central to building with large language models — tokenization, transformer architecture, prompt engineering, retrieval augmented generation, and production deployment — are thinly covered in computer science curricula, and existing interactive demos teach single concepts without a learning pathway. LLM Odyssey is an open-source, browser-based platform of 13 games arranged in three tiers aligned to Bloom's revised taxonomy: seven foundational Cognitive Core games, five Systems Forge games on production engineering, and Foundry Arena capstone challenges. Every game applies five pedagogical strategies drawn from the literature, including immediate formative feedback, scaffolded hints, difficulty progression informed by flow theory, worked examples, and scenarios taken from production practice. An initial Winter 2026 deployment at a Canadian college confirmed functional requirements and singled out adaptive difficulty as the top gap, and a mixed-methods evaluation protocol with 50 participants is specified for future study.
Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations
Defenses against insecure model-written code are all applied after the fact — static analyzers, fine-tuned classifiers, or a judge model reading finished output — ignoring what the generating model itself represents internally. The test here is narrower: whether the hidden activations at the last prefill token already encode whether the C or C++ code in context is vulnerable. Activations from Granite-4.1-8B, Qwen3.5-9B, Qwen3.6-27B, and Gemma-4-12B were used to train small multilayer perceptron probes, at 13.4 to 16.0 million parameters or under 0.2 percent of base model size, and evaluated on Devign, Big-Vul, Draper VDISC, and PrimeVul. On Devign the best probe reaches 68.8 percent F1, matching the 67.9 percent published fine-tuned-classifier state of the art while reading only a frozen general-purpose model's activations, though average F1 across all four benchmarks is 41.7 percent and probes trail substantially on the harder, more imbalanced sets.
FedPref: Federated Preference Learning for Structured Radiology Report Extraction
Radiology reports state findings and their locations in prose, while search and analysis need those relations in a fixed schema; the labels required to learn that extraction are spread unevenly across hospitals, and pooling patient data is often not permitted. FedPref has frozen public language models propose alternative JSON extractions, uses local annotations to rank them, and trains compact Qwen3-8B adapters federated across sites so only model updates are exchanged, with a heterogeneous pool of teacher models supplying contrast when repeated samples from one model collapse. Across six simulated hospitals with unequal data volume and disease prevalence it raises client-mean F1 by 2.49 points and worst-site F1 by 9.10 points versus isolated per-site training, with the biggest gains at the smallest sites; centralized training on the pooled preference pairs remains 2.66 points ahead, and the same ordering holds on a locked 400-report manually validated test set at 68.68 versus 71.67 F1.
Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss
De-identification tools miss protected health information (PHI) whose sensitivity is locally defined — hospital abbreviations, building names, internal codes — because those categories are institution-specific rather than pattern-matchable. On 100 annotated pediatric oncology notes from Texas Children's Hospital containing 5,322 PHI spans, eight large language models were run under three prompts of increasing specificity and compared against Stanford TiDE, OpenMed PII, and pattern baselines, plus 14 multi-agent and ensemble variants. The best model reached F1 0.918 against TiDE's 0.779, and no agentic architecture beat a single calibrated prompt; naming the missed categories recovered 79% of them, discouraging over-redaction restored precision, and model outputs exposed 414 candidate annotation gaps of which re-annotation confirmed 227 as real PHI the gold standard had missed.
Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting
Predicting which diagnosis codes a patient will receive at their next visit is prospective and multi-label — the target note does not exist yet and several codes can be right — and it pulls on two different model families: structured electronic health record foundation models that track recurrence and progression, and language models that generate flexible diagnostic hypotheses. ICD-Deepresearch composes both inside a deep-research loop, seeding candidates from SparseEHR, expanding them over two bounded research rounds with medical search and code dictionaries, adding an independent GPT-5 direct forecast, and jointly reranking under a fixed top-K budget before a separate module writes rationales. It reports patient-averaged precision/recall of 24.60/35.09% on MIMIC-III and 25.14/48.32% on MIMIC-IV, and physicians rated 51% and 68% of its retrieved documents useful versus 22% and 39% for standalone GPT-5 web search.
Structured Driving-State Narratives for Small Language Model-Based GNSS Spoofing Detection
Spoofed Global Navigation Satellite System (GNSS) signals can push an autonomous vehicle into a plausible but wrong position estimate, and detecting this means comparing what GNSS claims about vehicle behavior against what other onboard sensing says. The framework converts each independently derived driving state into a structured semantic narrative and hands it to a small language model (SLM) to flag spoofing and classify it among no attack, overshoot, stopped, turn-by-turn, and wrong-turn. Fine-tuned SLMs matched larger models on the same data, reaching 96.99% average accuracy with 97.18% F1 while needing lower inference latency and less GPU memory for both fine-tuning and inference, and the result held on unseen field data collected in Clemson, South Carolina.
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases
Legal case forecasting on judgments from the European Court of Human Rights (ECtHR) serves as a testbed for whether LLM reasoning is legally meaningful, with prompting strategies that vary in how explicitly they spell out what counts as sound legal analysis. Human and LLM-as-a-Judge assessments of GPT 5.4 find structurally complete but substantively shallow reasoning, and the expert-curated prompt yields more comprehensive analyses without improving prediction accuracy. The LLM judges were internally consistent but aligned only weakly with trained annotators, leading the authors to caution against relying on automated evaluation or treating task accuracy as a proxy for reasoning quality.
Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal
Reported accuracy on financial-news direction prediction depends heavily on whether the train-test split respects chronology, and an audit over 49,799 articles and 16 feature-model combinations — TF-IDF, MiniLM, FinBERT, fine-tuned RoBERTa-large and DeBERTa-v3-large, plus zero-shot, few-shot, and LoRA probes of Llama-3 and Qwen2.5 — finds that random splits inflate the Matthews correlation coefficient by 1.1x to 6.5x, with inflation tracking model capacity and feature richness. Conditioning on event type, mergers and acquisitions is the only category with a positive signal under chronological evaluation, and it fails to transfer to the FNSPID 2009-2020 U.S. corpus, localizing it to the study's 2024-2025 European-tilted data rather than a universal predictor. The authors liken chronological splitting to characteristics-purging in asset pricing — it strips the stale, predictable component of news — and call for leakage audits as required disclosure in financial NLP benchmarks.
Inductively Scalable, Single-Step Neural Surrogates for Wave-Scattering Inverse Problems
Neural surrogates promise electromagnetic simulation orders of magnitude faster than finite-difference time-domain solvers, but single-step, non-recurrent surrogates had only scaled to a few dozen controllable variables because random sampling of a huge configuration space wastes training signal. The fix is to generate training examples adversarially during training: a parallel process uses gradient ascent to hunt refractive-index and source configurations where the surrogate disagrees with a full-wave solver, stabilized by source and ground-truth normalization and an evolving replay dataset. The resulting two-dimensional scattering surrogate trains with up to 41,772 controllable variables and generalizes inductively to over 3 million without retraining, a 73.8x increase, matching or beating finite-difference designs for freeform beam splitters and gradient-index lenses up to 98 wavelengths wide at 1.29x to 26.5x speedups.
Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks
Detecting distributed denial-of-service traffic is hard when attack flows are a small, shifting minority of the data. GraphGAN turns sequential flows into k-nearest-neighbor graphs over sliding windows so feature similarity and temporal structure are both preserved, then uses a generator to synthesize realistic minority attack samples against a graph convolutional network discriminator, with a separate graph convolutional classifier trained on the rebalanced data making the final call. Across four benchmark intrusion datasets it reports higher accuracy, precision, and recall than prior methods, with the largest margins in data-scarce settings.
Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design
Designing biomolecules that bind each other from scratch is hardest for DNA and RNA, where complex-structure data is scarce and heterogeneous and the geometric and chemical constraints are sharper. MCTH (Monte Carlo Tree Hallucination) is an inference-only framework that treats pretrained folding and inverse-folding networks as frozen black boxes and uses Monte Carlo Tree Search to spend a fixed inference budget across competing all-atom sequence-structure design trajectories, steering on model confidence, uncertainty, and agreement or disagreement between multiple predictors, with optional biophysical control in the same loop. Across protein-RNA, protein-DNA, protein-protein, and protein-ligand design, adaptive search outperformed simpler sampling and cycling strategies at matched compute budget, and gains persisted under held-out evaluation with AlphaFold3 and Chai-1. No fine-tuning or backpropagation through the component models is needed.
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
Turning screenplays into shot plans for short-drama production depends on tacit directorial knowledge that is hard to write down, hard to evaluate against outcomes, and too bulky to inject wholesale into a prompt. SAGE (Skill with Attribution-Guided Evolution) derives episode-independent directing rules by contrasting each training screenplay with its expert storyboard, records which rules each narrative group adopted so localized feedback can be attributed back to individual rules and update them in place, and packages evolved rules into routed scenario bundles so each group retrieves only a bounded, situation-appropriate set. On 18 test episodes across three genres it scored 77.8 on an expert-validated rubric against 77.1 for professional directors, and across a 14-day deployment 87.2% of its 1,344 narrative-group outputs were accepted without substantive edits, cutting per-episode authoring time by over 83%. The authors release PROSE, 68 episodes pairing screenplays with professional-director storyboards.
Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task
Floating point operations (FLOPs) are the standard proxy for what it costs to run a large language model, but whether they track real energy consumption on software engineering tasks has not been tested. Using Morph, a many-objective optimization approach to knowledge distillation, the authors compare FLOPs-guided distillation against distillation guided by surrogate models that directly estimate CPU and GPU energy, on clone detection and vulnerability prediction, and extend the method to generative code summarization with CodeT5+. FLOPs proved an unreliable indicator of energy use, while energy-surrogate guidance produced student models that cut inference energy by up to 90% and memory by 86% for modest accuracy losses, putting them within reach of consumer hardware.
Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study
Independent audits of harmful content on video platforms are limited by the cost of video annotation and by moderation judgments that vary across languages. Sockpuppet accounts with four age personas (13, 16, 19, 40) in France, Italy, and Sweden collected 36,971 TikTok videos from passive For-You-page scrolling and from active sessions mixing scrolling with harm-keyword search, annotated by a multimodal LLM validated against native-speaker labels on a 300-video reference set — Gemini 2.5 Flash with eight sampled frames plus text scored best (aggregate kappa 0.42) at half the per-call cost of native-video upload, about $50 of API spend for a 10% sample. Keyword search returned 35-56% harmful content, a 1.5-7.5x increase over the scrolling baseline in ten of twelve country-age combinations, though the spike is temporary; under passive scrolling Italy had the highest harm rate at every age, peaking at 48.6% for the 19-year-old persona, and provider safety filters refused 1.1% of calls, under-counting the most explicit harms.
Benchmarking Automated Security Patch Backporting: How Far Are We?
Tools that automatically backport security patches to older code report success rates above 80%, but usually on datasets confined to one repository or version family. Porting Benchmark assembles 1,234 backporting cases spanning cross-version, cross-branch, and cross-repository scenarios and evaluates five tools — program analysis, LLM prompting, and LLM agents — under one common protocol. Aligned evaluation reshuffles the rankings, with FixMorph and Mystique degrading substantially while PortGPT and TSBPort hold up, and the best commit-level success rate falls from 85.2% on the simplest patch type to 24.0% on the most structurally complex. A 45-case subset with real tests and constructed proof-of-concept exploits further shows that reference-matching scores under-credit hard adaptations while missing integration failures that only execution catches.
GADR: Gathering Architecture Decision Records from Meeting Transcriptions
Existing LLM approaches to writing Architecture Decision Records (ADRs) assume reasonably structured input, whereas real architectural choices emerge from informal meetings where decisions are implicit, fragmented, and mixed with off-topic dialogue. GADR is a multi-agent, self-correcting workflow that extracts decisions from raw meeting transcriptions and drafts them in Nygard format. A feasibility study over five real project meetings, reviewed by four senior architects and evaluated by fifteen students, found the workflow captured most expert-identified decisions and beat zero-shot and few-shot prompting on stability and structural adherence, while retrieval-augmented enrichment deepened the drafts at the cost of content not traceable to the transcript.
TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification
A deployed text-to-SQL system has no ground-truth query or reference execution result to check against, so it must judge on its own whether a generated query answers the user's question; LLM judges and outcome reward models (ORMs) do this with little visibility into why. TraceSQL is a lightweight verifier built on 67 explicit diagnostic features covering question ambiguity and requirements, question-schema-SQL consistency, SQL structure, and intent alignment, so each prediction traces back to named evidence. On BIRD development databases it reaches 66.47% F1 and 64.48% ROC-AUC versus 61.87% and 58.26% for the GradeSQL-7B ORM baseline, with feature attribution showing it draws on both semantic grounding and deterministic SQL-structure signals.
ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction
Tabular foundation models such as TabPFN are effective but computationally heavy, and retraining them is often impractical, which makes few-shot in-context prompting an attractive alternative — except that choosing which training rows to use as examples is unsolved for tabular data. ARASH selects shots per query by analyzing the local neighborhood of that query within the training set rather than using a fixed example pool. The method cuts TabPFN prompt length by 1261.5× and memory usage by 2.56× at comparable accuracy.
Procedural Content Metageneration via Program Search and Continual Abstraction Discovery
Because language models emit executable programs, search can operate over procedural content generators rather than individual game levels; each run evolves complete Python generators through language-model mutation and crossover for Sokoban, Zelda, Dangerous Dave, and Lode Runner. Continual Abstraction Discovery (CAD) extracts reusable primitives from high-fitness programs into a run-specific helper module that later generations can call. A 2x2 experiment crossing CAD with access to a fixed hand-written domain API, comprising 160 complete runs of at least fifty generations each, finds that CAD raises mean final best fitness in all eight domain-and-API comparisons, with the learned libraries adopted by most later programs and repeatedly rediscovering validation, reachability, and structural utilities.
Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation
Radiology reports are dictated as free text, which complicates both downstream use and quality checking. A locally deployed pipeline of cooperating agents combined regular-expression rules with local large language models to sort report sentences into standardized anatomical sections and to flag quality problems: mismatches between Findings and Impression, gender-anatomy conflicts, and undocumented communication of critical results, applied to 638 chest, abdomen, and pelvis CT reports from 15 radiologists. The system restructured all 22,270 sentences and flagged 14.1 percent of reports, and in blinded review of a 45-report subset both radiologists agreed no clinically important information was dropped and no content was fabricated, with quality-assurance performance rated excellent or good in 84 percent of cases.
106 more specialized papers
- Ring-based Spatial Transformer: Learning Non-linear Spatial Interactions between Building Distribution and Pedestrian Flow Shun Nakayama, Takahiro Kanamori, Wanglin Yan
- An automatic-differentiation framework for time-lapse electrical resistivity tomography inversion of hydrologic dynamics Pu Yang, Zhengyang Fang, Yuxin Liu et al.
- Does the Heart Show Your Pain? Tackling the X-ITE Pain Challenge with Self-Supervised ECG Representation Learning Dominika Kunc, Przemys{\l}aw Kazienko, Stanis{\l}aw Saganowski
- In-Context Learning to Assess Built Environment Impacts on Perceived Neighborhood Walkability Among Mobility-impaired Older Adults Houhao Liang, Kresimir Friganovic, Joanne Kua et al.
- ARGUS: Attention-Guided Transformers for Scalable Person Identification Using Wi-Fi Telemetry Nayan Sanjay Bhatia, Pranay Kocheta, Yuhan Li et al.
- Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection Dominika Kunc, Przemys{\l}aw Kazienko, Stanis{\l}aw Saganowski
- Phase-Aware CNN for Real-Time 5G/6G Channel Estimation with Hardware-in-the-loop Validation Javad Zolfaghari-Bengar, Rakibul Rony, Elisa Gomez-de-Lope et al.
- RouteTS: Frequency-Time Routing for Time Series Forecasting Gaofeng Lin, Lei Duan
- Hardware-in-the-Loop Phase-Aware CNN for Real-Time 5G Channel Estimation Javad Zolfaghari-Bengar, Rakibul Rony, Elisa Gomez-de-Lope et al.
- Koopman early warning signals for bifurcation and rate-induced tipping Juan Nathaniel, Carla Roesch, Derek DeSantis et al.
- A Novel Fourier Feature Network for Solving Partial Differential Equations Qihong Yang, Zhijie Su, Yangtao Deng et al.
- Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks Kai Gu, Haizheng Zhong
- Generative Learning of Separatrices Ellis R. Crabtree, Dimitris G. Giovanis, Anastasia Georgiou et al.
- Iterative Refinement Diffusion for Super-Resolved Data Assimilation of Multiscale Physical Systems Mrigank Dhingra, Ramchandran Muthukumar, Rebecca Willett et al.
- The Note-Chord-Voice Framework: Structured Source Separation and Causal Inference for EV Charging Data Jiajie Chen, Jinfeng Li
- Real-Time State-of-Health Estimation and Online Degradation Prognosis from Partial Battery Discharge Using Physics-Informed Neural Networks Bego\~na Ispizua, Serio Gil-L\'opez, Leire Arrizabalaga et al.
- Uncertainty Identifies Difficult Samples Across Methods: A Multi-Task Study on a Heterogeneous Skin Lesion Dataset Leon Koole, Jiapan Guo, Matias Valdenegro-Toro
- NRCD: An Open Database of Collegiate Running with Unified Performance Standardization Jonathan A. Karr Jr., Ryan M. Fryer, Ben Darden et al.
- AMPLIFAI: A Multiphase CT Dataset for Benchmarking Clinical Reasoning in LI-RADS Assessment of Liver Lesions Pranav Kulkarni, Nikhil Shah, Amritansh Suryavanshi et al.
- A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification Christiaan M. Geldenhuys, Thomas R. Niesler
- M-LINKX: Multiview Graph Learning for Brain Cognitive Disease Detection An Phan, Yufei Jin, Xingquan Zhu
- ARISE: An adaptive residual-informed stability ensemble for feature selection in small-sample biomedical omics Zardad Khan, Amjad Ali, Naz Gul et al.
- When do machine-learned exchange-correlation improvements inherit into density-functional tight binding? Can Polat, Mustafa Kurban, Erchin Serpedin et al.
- Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task Alexandru-Stefan Morosanu, Valerian Cecan, Stefan-Daniel Achirei et al.
- Developing an Offshore Machine Learning Surface Layer Scheme Susan Dettling, Sue Ellen Haupt, Thomas Brummet et al.
- LLM-based Framework for Generating and Verifying Parallel DEVS Statecharts Vamsi Krishna Vasa, Hessam S. Sarjoughian, Edward J. Yellig
- A Physiology-Informed Digital Twin Framework for Simulating Liver Health Progression Sumaiya Afroz Mila, Sandip Ray
- Uncovering Hidden Leptonic Correlations with Flow Matching and Autoencoders Haruto Kitagawa, Satsuki Nishimura, Hajime Otsuka
- A Unified Mamba--MoE Surrogate for Closed-Loop Simulation and Measurement-Window Forecasting of Inverter Transients Haoguang Wang, Huy Hoang Le, Akhila Kandivalasa et al.
- Probability-Preserving Transformer for the Time-Dependent Schr\"odinger Equation Mushtaq Ali, Muzamil Tariq, Niaz Ali Khan
- Decision-Driven Regularization: A Blended Model for Learning and Optimization Gar Goei Loke, Qinshen Tang, Yangge Xiao et al.
- BrainLinear: A Linear Model for Brain Network Analysis in Sparse Tangent Subspaces Sijing Wu, Dongyuan Li, Miaoting Huang et al.
- Earth Observation Foundation Models for Terrestrial Ecohydrology: From Representation Learning to Process Inference Yi Yu, Jian Peng, Yucheng Lin et al.
- A Unified Geometric Framework for Developmental Analysis of Spatial Transcriptomic Data Mary Chriselda Antony Oliver, Kaitlyn Hohmeier, Tuyen Tran et al.
- SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences Tsz Fung Pang, Po Jen Chen, Nimish Ronghe et al.
- Detecting Money Laundering in Rwandan Mobile Money: A Machine Learning Framework Emmanuel Nahimana, Ya\'e Ulrich Gaba
- Rotation-Invariant Multi-IMU Activity Recognition under Independent Per-Location Orientation Shifts Seungyeol Baek, Yoonbyung Chai, Yonghyeon Lee et al.
- Sequential Multimodal Evidence Optimization for Product Media Ranking in E-Commerce Prasenjit Dey, Frank McIntyre, Arnab Sinha
- Integrating Persuasion Theory into the Epidemiological Modelling of Health Misinformation Spread on Social Media Mkululi Sikosana, Sean Maudsley-Barton, Oluwaseun Ajao
- BERTopic-Virality Prioritisation: A Scalable Framework for Thematic and Comparative Analysis of COVID-19 and Monkeypox Misinformation on Twitter Mkululi Sikosana, Sean Maudsley-Barton, Oluwaseun Ajao
- PLeDO: Pain Level Detection for Osteoarthritis from EMR Data Yuhao Chen, Jiahao Cai, Nafiz Sadman et al.
- Temporal Graph Prototype-conditioned Conformal Prediction for Fraud Detection Xudong Chen, Shengbo Gong, Lu Cheng et al.
- KOALA: Koopman Operator Learning for WiFi-Based Anticipatory Hum Quang-Anh N. D., Duc Pham Minh, Thao Phuong Pham et al.
- TransfHAR: Self-Supervised Wrist Representations for On-Demand Activity Recognition Aidan Bradshaw, Riku Arakawa, Xin Liu et al.
- Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes Farbod Abbasi, Zachary Patterson, Bilal Farooq
- Resource-Efficient QUBO Formulation for Anchored Currency Arbitrage Eric A. F. Reinhardt, Adam J. Hauser
- ReliaGate: Reliability Routing for Low-Stakes Wearable Stress Prediction Jaden Moon, Yu Wu, Arvind Pillai et al.
- Functional anatomy of Pythia-Herwig differences with Kolmogorov-Arnold networks Arghya Chattopadhyay
- Retrieval-guided Twin Fusion with Similarity-aware Contrast for Molecule-Text Alignment Shunshun Gu, Shengqi Qiu, Hang Zhou et al.
- Group ICA 2.0: Closing the Gap Between Subjects and Group Latent Decomposition with Copula-Linked Group ICA (CoLiG-ICA) Oktay Agcaoglu
- GOD: Enhancing Generalization via Deep Grafting for Sequential Recommendation WooJoo Kim, JunYoung Kim, JaeHyung Lim et al.
- Towards Reasonable Molecular Structure Elucidation from Infrared Spectroscopy with Chemical Feedback Yusen Tan, Hongyu Zhan, Hai-tao Yu et al.
- Representation Is Not Enough: Body-Localized Thermal Evidence for Contactless Stress and Craving Sensing in Opioid Use Disorder Sachin Deb, Harshit Sharma, Asif Salekin
- AsyTO: Asymmetric Temporal Operator for Parameter-Efficient Multivariate Time Series Forecasting Xiachong Lin, Du Yin, Hao Xue et al.
- RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction Mianzhi Liu, Fan Xiao, Zhiliang Yu et al.
- Multi-Feature Riemannian Hypergraph for Online Test-Time Adaptation of Motor Imagery Brain-Computer Interface Siqi Li (Peking University, Chinese Institute for Brain Research, Beijing) et al.
- Asymptotics-guided learning and symbolic regression for dispersive resonances Konstantinos Alexopoulos, Josselin Garnier
- Domain-Specific Text Embedding Models for Entity Resolution Khajesh Sapram, Srivardhani Raju, Kishore Konda
- RadioVIL: Anomaly-Aware Diffusion Models for Radio Map Inpainting and Zero-Shot Vehicle Localization Ruixin Zhao, Xiucheng Wang, Qiming Zhang et al.
- Quantifying the Gap Between Laboratory Battery Test Patterns and Field Duty Profiles Chunyang Zhao, Chresten Tr{\ae}holt
- Optimizing Multi-Market Participation of Battery and Electrolyser Systems Based on Field Performance Chunyang Zhao, Stoyan Trenchev, Shi You et al.
- Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic et al.
- Predicting, Evaluating, and Explaining Top Misinformation Spreaders via Archetypal User Behavior Enrico Verdolotti, Luca Luceri, Silvia Giordano
- Transfer Learning of Keystroke Dynamics for Cross-Device User Authentication Nuwan Kaluarachchi, Sevvandi Kandanaarachchi, Kristen Moore et al.
- FETERS: Few-Shot Early Time-Series Classification via Effective Ratio Selection Chen-An Tai, Yujia Wu, Vincent S. Tseng
- POI Recommendation with LLM-Augmented Multi-Graph Learning and Contrastive Alignment Burak Tamer, Wolfram H\"opken, Zehui Wang
- TRACE-CASH: Trial-History-Conditioned Reinforcement Learning for Adaptive Configuration Exploration in Time-Series CASH Yu-Han Huang, Yujia Wu, Vincent S. Tseng
- Self-Supervised Noise2Noise-Enhanced Denoising for Continuous-Scan Air-Plasma THz Spectroscopy Adam Umra, Oways Alsoloh, Oliver Nagy et al.
- Data-Driven Reconstruction of Spatially Resolved Electron and Ion Energy Distributions from Macroscopic Plasma Quantities with Deep Neural Networks Libin Varghese, Kaushik Prajapati, Bhaskar Chaudhury
- One Residual with Three Reuses: A Wristband Front End for Gesture Sensing Sam Rifaki
- Supervising the Path to Fine Scales: GalerkinFlow for Scientific-Field and Image Super-Resolution Zikang Zhan
- Learning Generalizable Reconstruction of High-Dimensional Neural Dynamics Anima Kujur, Zahra Monfared
- Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity Jiaqi Yao, Julia Kowal
- Turning spectra into images improves plant trait retrieval with 2D-CNNs Javier Lopatin, Teja Kattenborn, Eya Cherif et al.
- Non-Crossing Deep Quantile Regression for Distributional Survival Prediction Shuai Huang, Zhe Qu, Zhaowei Hua et al.
- Data-Efficient and Interpretable Classification of Circulating Tumor Cell Phenotypes in Microfluidic Devices via Deep Learning Serena Su, Yifan Wang, Senwei Liang
- A Data-Efficient Analytical Prior Machine Learning Framework for Sound Reduction Frequency Prediction in Helmholtz Resonators Jiaming Li
- Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval You Zuo (ALMAnaCH), Kim Gerdes (LISN, Qatent et al.
- CARA: Cognitive Adaptive Recommendation Agent Weijun Gao, Jinyang Dong, Chuanru Ren et al.
- Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence Jiaqi Wang, Huawen Hu, Shu Zhang
- AISA: AI Safety Assistant Framework for Continuous Improvement of Highway Construction Mason Smetana, Trevor Neece, Lev Khazanovich
- The Plot Thins: Uniformity and Linearity in Literary Summaries Rebecca M. M. Hicke, Sil Hamilton, David Mimno et al.
- Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection Chanwoo Park, Chanwoo Kim
- Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement Mohammad Talebi-Kalaleh, Qipei Mei
- Adaptive surrogate modeling for high-dimensional spatio-temporal output Berkcan Kapusuzoglu, Shunsaku Matsumoto, Yoshitomo Miyagi et al.
- NeuroAbs: A Neuro-Symbolic RTL Abstraction Framework for Property Checking Acceleration Zhiyuan Yan, Xiaofeng Zhou, Ziyue Zheng et al.
- LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap Yining Hua, Cyrus Ayubcha, Hongbin Na et al.
- MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting Bowen Liu, Mingming Sun
- ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generation Weiran Wang, Hongxiang Shi, Huitao Tang et al.
- From Entity Mentions to Tone: An LLM-Based Pipeline for Media Bias Analysis Klesti Hoxha, Olti Qirici
- Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery Mohammad Javad Ahmadi, Hamid D. Taghirad
- CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method Jin Su, Zhuofeng Zhao, Huanhuan Wang et al.
- Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff
- DMT-Dens: Density-preserving manifold visualization for biological data Ruizhe Wang, Yixuan Dong, Bolin Yang et al.
- From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support Ngoc Luyen Le, Marie-H\'el\`ene Abel, Bertrand Laforge
- Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models Sahab Zandi, Noah Kostesku, Christophe Mues et al.
- Learnware for CSI Feedback: Scene-specific Small Models Can Do Big Xiangyi Li, Jiajia Guo, Chao-Kai Wen et al.
- MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil et al.
- Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks Matin Amoozadeh, Amin Alipour
- SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE Xuan Zheng, Kento Uchida, Shinichi Shirakawa
- Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection Bin Li, Dongdong Wang, Siyang Lu
- When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era Lotta Kiefer, Brisca Balthes, Christoph Leiter et al.
- Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media Yijie Xu, Chao Wang, Hui Xiong
- Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach Lu Xu, Xu Li, Linjiang Zheng et al.
- Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors Mahdi Saberi, Ya\c{s}ar Utku Al\c{c}alar, Merve G\"{u}lle et al.
- HLSR: Hybrid Live Forecast Selective Dynamic Vehicle Rerouting for Real-Time Congestion Avoidance Xiao Wang, Shun Ren Yang, Hui Nien Hung
Large Language Models 78
When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation
Five uncertainty estimators — mean token entropy, verbalized confidence, P(True), entropy ensembles, and semantic entropy probes — are tested on three small code language models across HumanEval and BigCodeBench to see whether uncertainty methods built for natural language transfer to code. Multi-sample P(True) correlates best with correctness while the rest correlate only weakly, and routing those signals into self-correction backfires: uncertainty-driven regeneration lowered Pass@1 in 5 of 6 configurations, by 3 to 10 percentage points, with adaptive decoding hurting in 4 of 6. Only execution-verification-based regeneration reliably helped, adding 6 to 26 points on HumanEval and 8 to 20 on BigCodeBench, with the largest gains for the weakest baselines. The conclusion drawn is that cheap uncertainty scores are useful as gates on costlier verification loops rather than as substitutes for them.
Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation
Language-model judges scoring responses against fine-grained rubric checklists usually evaluate each rubric in a separate inference call; batching them into one pass is cheaper but introduces rubric interference, where the verdict on one rubric shifts depending on which others are present. A measurement framework probes this through rubric set expansion, subsetting, reordering, and noise injection, and finds only about one-third of samples receive fully consistent verdicts across differing rubric compositions. Self-Anchored Rubric Alignment (SARA) treats the model's own single-rubric judgments as stable anchors and aligns its multi-rubric reasoning to them via on-policy self-distillation, needing no external supervision. Consistency improves on HealthBench, FLASK, and ResearchQA for Qwen3 and Llama-3.1 while agreement with base models and with GPT-4.1 as reference judge holds, and the gain transfers across datasets.
Rethinking Reverse KL as Adaptive Entropy Distillation
Distillation objectives for compressing large language models typically blend forward and reverse Kullback-Leibler divergence, treating the reverse term as one fixed half of a mixture. Decomposing on-policy reverse KL into a teacher-fitting term and a student-entropy term shows that the token-level optimal student is a tempered version of the teacher, with an adaptive weight setting the trade-off between mode-seeking and uncertainty preservation — a control already latent in reverse KL, requiring no explicit forward-KL branch. Adaptive Entropy Distillation (AED) uses the teacher's entropy to calibrate per-token imitation strength, reporting better overall results on instruction-following and mathematical reasoning benchmarks alongside closer student-teacher distributional and entropy alignment.
The Quantum Shortcut: Complex Phase-State Dynamics Reduce the Optimization Steps of Sequence Models
Sequence models are usually distinguished by backbone — attention versus recurrence — but the substrate underneath, a real-valued hidden state with an affine-softmax readout, is nearly universal and rarely varied. Swapping in a complex-valued alternative from quantum mathematics, where information rides in the state's phases and scores are quadratic Born forms, and relaxing the two blockers of exact unitarity and Born vocabulary readout, the substrate is instantiated in both the Mamba state-space model and a Transformer at 253M parameters matched to within 0.02 percent. Under one fixed protocol on three byte-level corpora, the complex models reach every measured validation loss in roughly one third (state-space) and one half (attention) of the optimization steps of their real-valued counterparts. The backbones then diverge: past learning-rate warmup the state-space advantage keeps widening, reaching 0.354 bits per character on OpenWebText and 0.396 on FineWeb, while the attention advantage decays toward zero on every corpus, marking it an early-training effect.
Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data
Each row of a transformer attention matrix is a probability distribution, and in trained models most mass lands on a "sink" token, so comparisons using cosine similarity, Jensen–Shannon divergence, or entropy depend on an unreported choice: keep the sink or drop it and renormalize. Across ten pretrained models from five families, 17–47% of verdicts about which of two heads is more similar flip with the convention, and the dominant structure in a standard BERT head-clustering pipeline turns out to be an artifact of it. Treating rows as compositional data separates the two mixed questions exactly, with Aitchison distance splitting orthogonally into a sink term and a content term and entropy splitting by an exact identity. Under that split, most measured entropy collapse during training is the sink growing rather than attention sharpening (95% of the drop at 1B parameters), and pruning heads through the wrong channel can inflate perplexity more than a hundredfold.
Tail-Aware Top-$k$ On-Policy Distillation
On-policy distillation trains a student language model to match a teacher's next-token distribution along the student's own trajectories, usually by minimizing reverse Kullback-Leibler (KL) divergence over the teacher's top-k tokens after renormalizing them. That renormalization discards the total probability sitting outside the top-k, which lets the student's tail probability and entropy creep upward and empirically hurts downstream accuracy. TA-OPD restores the missing signal by adding a single extra token that carries the tail mass to the top-k reverse-KL objective, keeping the student's distribution aligned with the teacher's without extra cost. Reported gains reach up to 8.05 points on Avg@8 across common benchmarks.
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
When a user's question is underspecified, a model should notice, identify what is missing, ask for it, and only then answer — a behavior formalized here as solving a k-underspecified constraint satisfaction problem, where k counts the variables jointly needed to pin down the target. MT-InfoSeek instantiates this with 5,251 problems and 9,006 task instances across mathematics, logic, biology, medicine, and general knowledge, scoring what models ask, when they ask it, and how acquired information changes the answer. Models recognize that more information is needed but systematically underestimate how much — on logical problems at k=2 they under-predict the degree of missing information about four times as often as they over-predict it — and they fail to find a minimal sufficient query set, improve only marginally when told the true k, and often stop asking too early. Measuring final sufficiency separately from answer accuracy exposes model differences that accuracy alone hides, supporting the claim that multi-turn information seeking is a distinct capability current evaluations do not capture.
Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification
Open-weight checkpoints get fine-tuned, quantized, pruned and merged with no reliable record of what descended from what, raising the question of whether weights alone reveal shared ancestry without any data or model queries. Residual training leaves a shared identity-aligned component in branch products that would fool a naive comparison, so the method removes it and compares the remaining checkpoint-specific structure across residual blocks, producing a symmetric lineage score calibrated against independent checkpoints. On residual-MLP and GPT-2 benchmarks the score perfectly separates fine-tuned, LoRA-merged, pruned and quantized descendants from independent and distilled models (AUROC 1.0), cleanly distinguishing weight ancestry from mere behavioral similarity, and it survives function-preserving "checkpoint laundering" that degrades or breaks weight-space baselines while running 76x faster than the nearest robust one. The signal appears across six model families, and a case study correctly classifies 3 related and 7 unrelated public LLaMA-2 checkpoints.
T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework
Language models can propose high-level code transformations for optimization but cannot themselves establish that a transformation preserves semantics, which limits how aggressively they can be applied. The T-LLM Compiler closes that loop by combining model-generated transformations with traditional compilers and formal verification tools, feeding verification failures back as corrective iterations. On PolyBench/C benchmarks it reaches up to 83.3% optimization accuracy with transformed code averaging 26.7% speedup over standard baselines, and the source is released openly.
Certifying Compressed Language Models: An Audit and a Statistical Toolkit
Claims that a compressed model is equivalent to its original usually rest on a fraction of a point of benchmark accuracy, a number that hides per-item changes in opposite directions cancelling each other out. Mining 1,707 paired model-by-task cells from public per-item evaluation dumps spanning 1.3B to 405B parameters, the authors find that churn, the rate at which the two models disagree on individual items, runs roughly five times the net accuracy delta, with cells that score identically still disagreeing item by item. A preregistered audit of 17 equivalence claims from method papers, model cards, and vendor documentation found none declaring a prospective numerical margin and none releasing task-matched per-item outputs. They supply paired equivalence testing at a declared margin with certification tables for required sample size, and show in a controlled GPTQ versus AWQ comparison that simply changing the calibration draw reversed the observed method ordering in 5 of 8 confirmatory cells.
RecurrentGPT: Expressive Depth through Recurrent Modulation in Transformers
Giving each transformer layer its own weights preserves functional specialization but costs memory, while naive weight sharing across depth collapses that diversity and hurts quality. RecurrentGPT brackets a single shared core block, iterated R times, between fixed prelude and coda blocks, and modulates each iteration with a lightweight projection and an elementwise update gate conditioned on the hidden state, the prelude output, and noise resampled at every step. Under matched FLOPs a 3-layer version matches a 12-layer GPT-2 Small and leads MoR and heavy-tail depth sampling in all nine scale-by-budget cells, and at matched parameters and data it reaches 2.76 validation loss versus 2.84 without recurrence. At large scale the trade is 63% fewer parameters and 59% less peak decoding memory for a 10% increase in generation latency.
The Distributional View of Knowledge Distillation
Token-level knowledge distillation compares teacher and student distributions pointwise, so a Kullback-Leibler gradient cannot distinguish which wrong token absorbs the misplaced probability mass. The proposed distributional view represents the teacher as a family of multi-temperature views along the annealing path of its logits and trains the student against a geometry-aware aggregate under an embedding-based ground cost, formalizing mixtures, log-linear pooling, entropic Wasserstein barycenters and a debiased Sinkhorn-divergence variant, and proving that log-linear pooling of tempered views collapses exactly to a single temperature. Experiments on instruction-tuned Pythia pairs yield three empirical laws, the sharpest being that which distillation loss wins is governed by the ceiling gap between supervised-fine-tuning and teacher perplexity rather than being a fixed property of the loss — when the teacher barely beats a supervised student, no distillation beats plain fine-tuning, and the ranking of losses inverts once a real ceiling exists.
MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation
Sparsely-activated Mixture-of-Experts (MoE) Transformers route the same number of experts at every layer, ignoring that layers differ substantially in redundancy. MAPLE reallocates a fixed routed-expert budget across layers of any pretrained MoE model without touching weights or retraining, probing each layer's sensitivity to expert count, deriving a closed-form optimal budget assignment, and refining it with a sensitivity-guided genetic search. On DeepSeek-MoE-16B it uses only 75% of the experts yet beats the full 100% uniform baseline on ARC-E, ARC-C and BoolQ (65.09 to 71.40 on ARC-E), and implemented in SGLang it cuts single-GPU end-to-end serving latency by 32.2% and raises throughput by 47.4%.
Spectral Rank Certification for Foundation Model Adapters
The rank you pick when training a LoRA adapter is a design choice, not evidence about how much structure the adapter actually carries. This work builds a finite-sample framework for inferring effective rank from adapter spectra, centered on an exact chi-square divergence for the Gaussian rank-one reference experiment with the signal direction integrated under a rotation-invariant prior, yielding a computable Le Cam bound at realistic layer sizes plus the rectangular Baik-Ben Arous-Péché limit, and an empirical-null workflow of factor reconstruction, Monte Carlo p-values, stagewise and block testing, and Benjamini-Hochberg reporting. Auditing 26 public adapters covering 684 modules across six architecture families and 31,770 spectra rows, the authors find that calibrated effective rank is usually far below nominal rank and diverges systematically from the common 95% energy-retention heuristic.
SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning
Parameter-efficient fine-tuning (PEFT) methods that share weights across layers either share uniformly, which slows convergence, or use dynamic masking, which costs extra compute. SAPE (Sandwich Adapters for Parameter Efficiency) instead uses a fixed "sandwich" topology: intermediate Transformer layers are routed through balanced shared group adapters while the input embedding and final projection boundary layers get isolated adapters to avoid gradient interference. It beats proPETL on RoBERTa-large with 10% of the parameter budget, and under a ~0.6M-parameter cap on LLaMA-3.2 (3B) it improves on AdaLoRA by +4.85% on GSM8K and +3.11% on CommonsenseQA; ablations show hard sharing regularizes semantic generalization but slightly blunts the sharp layer-wise transformations arithmetic reasoning needs.
FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA
Federated fine-tuning with Low-Rank Adaptation (LoRA) faces a dilemma: averaging the low-rank factors separately preserves each client's optimization continuity but aggregates the update inaccurately, while reconstructing in the product space fixes the aggregation error yet hands clients freshly factorized weights that discard their local state. FedPA-LoRA keeps each client's own factors across rounds and aligns their product toward a rank-specific global reference, while the server aggregates heterogeneous-rank updates in the product space and reconstructs a rank-constrained global adapter without materializing the dense matrix, with convergence proved for both equal and unequal client ranks. On language understanding and generation tasks it beats the baselines across heterogeneity levels, gaining up to 6.82 percentage points of average GLUE accuracy under heterogeneous client ranks.
Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference
Sparse mixture-of-experts (MoE) models activate few experts per token but still need the whole expert bank resident in memory, which is the binding constraint on small GPUs. ExactMoE quantizes only the routed experts to symmetric group-128 four-bit weights, stores them in kernel-native MARLIN layout in pinned host memory, and streams selected experts through a GPU-resident slot cache with fused grouped kernels, leaving the router, attention, embeddings, and language-model head in BF16 and keeping top-k routing and full expert availability unchanged. On OLMoE-1B-7B-0924-Instruct on a single NVIDIA L4, a 16-slot configuration cuts peak reserved GPU memory from 14.168 to 1.836 GiB (87%) while keeping 81.85% of BF16 decode throughput, a fully resident 64-slot setup actually outruns BF16 at 31.9 versus 21.7 tokens/s, and accuracy across 12,450 zero-shot multiple-choice questions holds at 99.23% of baseline.
Language models suffer from a curse of ambiguity
As language models increasingly bootstrap their own improvement by sampling from themselves, how faithfully they reproduce the true next-token distribution matters more, yet some distributions are intrinsically harder to fit. The authors identify a "curse of ambiguity" affecting language models and, more broadly, any neural network emitting discrete distributions: the more ambiguous a next-token distribution is, the less accurately it is learned, because ambiguous distributions need more capacity to store, larger embeddings to represent, and more optimization steps to fit, while also amplifying token-sampling noise. Theory is validated on synthetic tasks with known ground truth, and the same signatures appear in models trained on real text, yielding a practical rule for when a model's output distribution can be trusted.
Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off
A decade of attention mechanism research is surveyed along four threads — sequence-to-sequence formulations, adaptation to vision, efficiency work attacking the quadratic cost, and interpretability — covering Bahdanau/Luong alignment, the Transformer, sparse and linear attention, IO-aware exact methods such as FlashAttention, state-space alternatives including Mamba, plus induction heads, superposition, and attention-state-space duality. Twenty-one methods are scored on efficiency, expressiveness, and interpretability (EEI) by a single rater, and a 200,000-sample Monte Carlo analysis with an assumed ±1-point perturbation finds rank changes of more than one position in 67-70% of samples, with a rank-matched null model showing a similar profile, so the authors present the scores as supporting only coarse tier comparisons. The survey adds a benchmark synthesis with cross-study caveats, a five-problem gap analysis, and a 2015-2026 timeline.
Do Language Models Consistently Encode the Current Year?
Temporal reasoning presumes a model knows what year it is, so two probes test that belief separately: an associative task that infers the year from verb tense, and a declarative task that just asks. Both land within a year of the instruction-tuned models' post-training cutoff, and for base models the associative probe estimates the pretraining cutoff with an average error of only 10 months across 13 models — but the mechanisms differ, with the associative year running through factual-recall-like circuits and the declarative year lacking consistent causal pathways. That split makes the current year hard to update: prompting shifts the declarative year in 94.6% of 351 target years but the associative year in only 1.7%, year-shifted supervised fine-tuning moves the associative year in just one of eight models, and weight editing works on each task alone without generalizing across both.
DeltaLog: Deferred Materialization of Recurrent States for Linear Attention Decoding
Linear attention replaces the growing key-value cache of softmax attention with a fixed-size recurrent state, but decoding implementations typically write the entire state back to memory after every token, so state maintenance dominates memory traffic for models with large states and many heads. DeltaLog keeps a dense base state plus a bounded log of recent compact update factors: most decode steps only append to the log, and periodic merges fold accumulated updates back into the base, so the model sees exactly the same dense state as eager decoding. Implemented for GDN, KDA, and RWKV6 inside a serving stack, it speeds the state-update kernel by up to 1.86x, cuts profiled recurrent-state write traffic by up to 7.83x, and yields 1.05–1.20x end-to-end serving gains.
SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization
Weight-only post-training quantization (PTQ) shrinks large language models to fit tight memory budgets, but accuracy tends to collapse at 2-3 bits because backpropagation-free optimizers make per-group decisions without accounting for how the remaining continuous weights could compensate, and because they leave the affine quantization grid fixed. SchurOpt analytically eliminates the optimal continuous response of the suffix, producing an exact groupwise quadratic with Schur-complement curvature, then alternates closed-form row-wise scale and zero-point refitting with coordinate descent over the integer codes; holding the GPTQ objective fixed, this alone adds 11.88 percentage points of mean zero-shot accuracy on 2-bit Qwen3-4B. Because tighter reconstruction stops helping at higher precision, the full SchurQuant adds quantized-prefix teacher reconstruction, reference-weight regularization, residual-add targets, and teacher-decision token weighting, and across eight Llama and Qwen models beats the strongest backpropagation-free PTQ baseline by 9.65 percentage points at 2 bits.
GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix
Multi-agent serving workloads have two key-value cache regions with opposite needs — a long shared prefix that wants contiguous storage, and per-request suffixes that want fine-grained allocation — yet production paged engines apply one page size to both. GraniKV splits them, putting the shared prefix in a contiguous HOT pool and suffixes in a token-level COLD pool, with a per-step dispatcher picking between two attention backends depending on whether the step is compute-, memory-, or communication-bound. At 16K shared prefix tokens it reaches 2.16x, 1.98x, and 1.57x output-token throughput over the production baseline on Llama-3.1-8B, Qwen-2.5-14B, and Qwen-2.5-32B; most of that comes from cascade attention at saturation, but under heterogeneous serving with distinct prompts of differing lengths the storage layer alone sustains 1.95x while batch-global cascade drops to parity.
FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy
Binary weight quantization promises extreme compression for large language models, but the gains often evaporate in practice because implementations still lean on floating-point arithmetic or dequantize at runtime. FluxBin co-designs the algorithm and the kernel: on the algorithm side, Decoupled Row-Column Binary Decomposition raises representational capacity while staying hardware-friendly, and Hessian-guided saliency selection keeps a hybrid of bases for the most important weights; on the kernel side, a CUDA lookup-table construction with fused scales removes most floating-point work, and Virtual Columnar Mapping packs the irregular sparse salient matrices into dense execution. Reported results reach up to 5.92x speedup and 10.19x energy savings at accuracy comparable to heavily fine-tuned methods, with 4x memory reduction letting a 70B-scale model run on a single A100.
SubZero+: Efficient Zeroth-Order LLM Fine-Tuning via Large Learning Rates
Zeroth-order optimization fine-tunes language models without backpropagation, but its gradient estimates carry high variance, making convergence fragile and very sensitive to the learning rate. SubZero+ stabilizes the SubZero framework three ways: multi-query gradient estimation inside layer-specific low-rank subspaces that avoids the usual multi-query paradox, an Adam variant that adapts using in-subspace gradient statistics, and a sign correction for QR-based subspace construction so projection matrices are properly Haar-distributed. On models from 1.3B to 32B parameters evaluated on SuperGLUE under both full-parameter tuning and LoRA, it beats prior zeroth-order baselines, widens the range of learning rates that remain stable, and narrows the gap to first-order training at minimal extra memory cost.
Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
Compression is what makes large models fit on edge hardware, but the paper's organizing point is that what compresses well does not necessarily deploy well. Dozens of recent works reporting results on real hardware are surveyed into practical guidelines, which are then followed to deploy compact language and image models on GPU, CPU, and Raspberry Pi for question answering and image segmentation. No single technique wins: Qwen3.5 0.8B reaches 93.85 SQuAD F1 under Q5_K_M GGUF quantization while structured pruning at the same precision costs 16 F1, yet for segmentation the ranking flips, with pruning cutting size by nearly 80% at near-constant mIoU. Pruning can even inflate the deployed artifact by 21-49% by breaking k-quant super-block alignment, pushing Raspberry Pi latency up 3.4x, and one LoRA-recovered variant holds 71% strict BoolQ accuracy while sending 97 of 100 predictions to a single class — compression manufacturing apparent competence rather than visibly destroying it.
Global Simulation-Guided Dynamic Operator Scheduling for Efficient Multi-Tenant Model Serving
Scheduling GPU work at container granularity in multi-tenant model serving leaves many short idle slices inside containers unused, while moving containers around is too slow to exploit them without breaking service-level agreements. SliceScheduler instead schedules individual operators, using a Global Mapping Graph that tracks operator dependencies, tensor shapes, resource mappings and live execution state across the cluster, plus a simulator that predicts execution and memory evolution under candidate placements before committing to them. Implemented as a PyTorch backend and evaluated on production trace replay, it improves token throughput by 1.10 to 2.29 times over existing approaches while keeping service-level-agreement violations within 9%.
Routing Divergence Is Not Evidence of Behavioral Influence in Same-Weight MoE Self-Distillation
Two Mixture-of-Experts forward passes over the same frozen weights can route a token to different experts, which raises the worry that self-distillation setups where a demonstration-conditioned teacher supervises a query-only student are influenced by routing changes rather than content. An exact blockwise decomposition splits the discrepancy into a routing term, which moves gates at fixed content, and a dense-like content term, measured across seven open-weight checkpoints and two domains. The routing term moves outputs by less than half the natural context effect and is largely reproduced by matched-norm noise, whereas the content term is strongly direction-specific, so router divergence on its own is not evidence of behavioral influence; the authors recommend measuring residual-stream exposure first and running a behavioral intervention when the answer matters.
The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT
Encoder-decoder speech recognition (ASR) and neural machine translation (NMT) systems will happily emit fluent text when the input contains no recoverable message at all. The question asked here is whether the score assigned to the reserved null token — the model's option to stop generating — already encodes a usable abstention signal, audited across several recognizers and translation models via native null-token scores and scalar logit shifts, with additional decoder-state probes, supervised row edits, and comparisons to external gates for Whisper. The signal is usually present but stock decoding does not act on it, and while raising the null-token score sharply suppresses fabrication, aggressive intervention deletes valid speech and truncates legitimate translations — so abstention methods should be scored on deletion cost alongside hallucination reduction.
Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation
Language models make capable conversational recommender systems, but steering a multi-turn dialogue to actually extract a user's preferences is unsolved: existing methods bolt on a separate reinforcement-learning agent with templated questions, or optimize for interactivity as judged by another model, without ever measuring how much information a turn gained. The proposed reward scores each interaction by how much it reduces the assistant's own uncertainty, computed as entropy over its candidate recommendations, which requires no ground-truth recommendation labels and so applies where they are unavailable. Fine-tuning with this signal under both supervised fine-tuning (SFT) and direct preference optimization (DPO) improves recommendation quality and conversational efficiency on the INSPIRED and ReDial datasets.
A Scalable Pipeline for LLM-Teacher Distillation Labeling: Work-Stealing Job Scheduling and Memory-Aware GPU Concurrency
Labeling millions of text items with a teacher model raises two practical questions — what label quality each dollar of teacher buys, and how to keep GPU workers busy under skewed, crash-prone workloads. The pipeline answers both with a work-stealing ring pool where each worker drains its own queue then steals from ring successors, using atomic conditional writes for exactly-once task claims and stale-claim sweeping for crash tolerance (implemented over a single SQLite file, so the reference version has no dependencies); a memory-aware rule that sizes per-node parallelism by how many model copies fit on the GPU; and a relabeling benchmark where the teacher relabels a gold-labeled public dataset so quality becomes an agreement measurement and cost falls out of measured throughput. Under skewed load the pool sustains up to 3.4x the throughput of static sharding and matches it at zero skew, and when half the workers are killed mid-run it loses 0 of 2,000 tasks against static sharding's 953; all experiments run on public data and commodity hardware with code, tests, and logs released.
Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards
Preference benchmarks depend on hired annotators whose identity is normally treated as an implementation detail, and measuring that detail shows it matters a lot at the item level: on the 2,885 MultiPref items where both pools are internally unanimous, so no tie-breaking rule applies, expert and crowd annotators assign different majority labels on 23.6% and name opposite winners on 9.2%, with comparable figures of 30.5% and 8.5% on 246 unanimous MT-Bench cells. The resulting six-model leaderboards are nonetheless bit-identical (Kendall tau = 1.00) — but that invariance is shown to be weak evidence, since pool switching moves win rates by 1.9 percentage points, one adjacent pair had a 38% chance of swapping, an item-level bootstrap displaces a model in 28% of resamples, and the same perturbation displaces a ten-model leaderboard with probability 0.86 and a twenty-model one with probability 0.9997. The work also shows a widely used dataset's stated assumption of no intra-group annotator variability is false, and that an LLM judge tracks the crowd pool over the expert pool on all three models tested, including one from a different vendor.
Conditional Evaluation of Language Models with Cheap Auxiliary Signals
Aggregate benchmark accuracy conceals which inputs a model handles well, and estimating conditional performance from gold labels alone is expensive, while cheap signals like LLM-judge scores, pairwise comparisons, and confidence are available for every item but biased. LACE (Local Augmented Control-Variate Evaluation) applies local centering — subtracting the cheap signal's conditional mean within the target profile region — so any linear augmentation has zero conditional mean and cannot shift the estimand, leaving the augmentation coefficient to affect only variance. The authors prove calibration-free identification, unbiasedness for grouped profiles, local oracle optimality among centered linear augmentations, and first-order adaptivity, with efficiency gains governed by a local R-squared; the estimator is tested on MATH-500, ScienceQA, MMLU, WinoGrande, HellaSwag, TruthfulQA, GSM8K, and ARC.
Beyond Binary Priorities: Multi-Tier SLA Scheduling for Large Language Model Serving
Production LLM serving mixes latency-critical API traffic with background batch jobs, but Llumnix, a migration-capable multi-instance inference scheduler, supports only high and normal priority — too coarse for real service-level agreement tiers. This work generalizes that priority model to arbitrarily many tiers with per-tier headroom under exponential decay and tier-aware dispatch ordering, implemented alongside the full migration pipeline inside the Vidur inference simulator and compared against INFaaS, vLLM, Orca, and Sarathi-Serve under uniform, Gaussian, and enterprise priority distributions. Four priority tiers give the best cost-effectiveness, with prefill mean speedups up to 8.3x and end-to-end P99 speedups up to 3.1x over INFaaS and 46-68% better cost-per-latency, and the scheduler holds up at ten tiers without tail latency collapse.
Architecture-Dependent Causal Transfer of Activation States Across Large Language Models
AI systems currently exchange information as natural language text, paying encoding, token, and latency costs, which raises the question of whether internal activations could be piped between models instead. Using Qwen2-0.5B, Phi-3-mini, Mistral-7B, and FLAN-T5-base, the authors measure representational alignment, train a projection network for cross-model retrieval, and test end-to-end causal transfer by injecting projected activations during generation. Retrieval works well above chance for the three decoder-only pairs (45-50% top-1 versus 5% chance) but at chance for the encoder-based FLAN-T5, and injection produced a statistically significant causal effect for only one of three decoder-only pairs, which the authors read as transfer of the representational vehicle rather than of meaning — making activation-state transfer architecture-dependent rather than universal.
Evolving Executable Pipeline Programs for AutoML with Language Models
Automated machine learning systems can only assemble pipelines from a preprocessor-learner-hyperparameter space fixed in advance, never inventing structure outside it. LACE instead evolves complete executable programs: a population of scikit-learn-compatible Python classes is mutated and recombined by a large language model acting as the variation operator, under a leakage-controlled protocol that hides dataset identity from the generator. On 68 OpenML classification tasks, LACE driven by GPT-5.4-mini significantly beats auto-sklearn, H2O, and a fixed XGBoost baseline with no detectable gap against AutoGluon, and because candidates are ordinary Python, the returned pipeline is directly readable and editable, with the component set extended by changing the prompt rather than the framework.
Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN
Serving large language models from cellular base stations means a mobile handover can strand a request's key-value (KV) cache at the source station: keeping inference there inflates inter-token latency indefinitely, while rebuilding at the target only after handover prolongs the service interruption. Pallas starts preparation before the handover fires, splitting the token sequence into a stable historical prefix that the predicted target recomputes locally and an evolving suffix whose KV blocks the source streams across, with an online scheduler choosing how early to begin from mobility predictions and runtime telemetry. A vLLM-based prototype across three models and 100–500 Mbps inter-station links cuts average service interruption time by 2.28× to 89.68× versus target-side recovery and lowers inter-token latency 16–50% versus forwarding from the source.
Toward Better Assessment of LLMs' Performance in Clinical Error Detection
Clinical error-detection benchmarks are built by injecting errors into notes, so each erroneous note has a clean counterpart, yet aggregate metrics like F1 and balanced accuracy score notes in isolation and ignore that pairing. Evaluating 15 large language models on 4 standardized test sets in 3 languages, 13 of 15 fall below random pairwise discrimination despite F1 scores that would ordinarily read as moderate, with bias flipping by language — the same model defaults to 'no error' in one and over-flags in another. A procedure for scoring cited evidence shows models reliably locate error-relevant content but then fail to give the correct verdict on the clean counterpart, and because F1 and pairwise accuracy move in opposite directions under the same bias, F1-based leaderboards can promote the weakest discriminators.
Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
Interpretability and chain-of-thought faithfulness research both produce explanations of model behavior, but rarely test them by the standard of counterfactual simulatability — whether the explanation lets you predict what the model does on related, edited inputs. CHIVE (Counterfactual Hypothesis Investigation Via Edits) is an agentic pipeline that spots unexpected model behaviors in real usage and probes them with counterfactual prompt edits, producing thousands of explanations paired with supporting evidence. Applied as an evaluation, it finds no uplift from any of the common large language model interpretability techniques studied in predicting counterfactual behavior; applied as a data generator, training models to predict the outcomes of its counterfactual experiments transfers to a range of out-of-distribution settings.
Proteus: Incremental Memory Activation for Long-Context Sequence Modeling
Memory-based sequence models compress long contexts into a compact state to escape attention's quadratic cost, but they expose the same memory capacity from the first token onward, so early tokens face no compression pressure and crowd out later context. The incremental memory activation paradigm instead expands the memory's effective capacity as the sequence grows: an early bottleneck forces harder compression of history, and newly unlocked capacity absorbs later context with less interference. Proteus implements this at no extra cost and drops into a broad class of memory architectures, including SWLA, Comba, Titans, and Hope-Attention, giving consistent gains on language modeling, reasoning, and long-context retrieval that grow with context length.
A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications
Security research leans on expert surveys, but recruiting Security Operations Centre (SOC) analysts is hard given workload, burnout, and confidentiality, leaving small samples that large language models might cheaply supplement with synthetic respondents. Real SOC professional responses are used as ground truth to compare persona-based and aggregate LLM-generated answers across multiple models and prompting settings, measuring stability, inter-model agreement, and alignment with humans. Answers are internally consistent yet systematically diverge from experts, showing reduced variance, central tendency bias, and homogenised opinions, supporting LLM respondents for piloting and hypothesis generation but not as replacements for expert elicitation.
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
What an API buyer purchases is a dated contract rather than a model name: it bundles the requested and served model, the reasoning-effort setting or its absence, output rail, prompt, and price schedule. A preregistered paired contrast ran Sonnet 5 with explicit high reasoning effort against the same model with the effort term omitted, over 30 AIME 2026 problems at five calls each, with every paid attempt assigned one frozen terminal category and confidence intervals from resampling items. Explicit high effort cost $0.01031 more per call on average, while the accuracy difference of +0.0133 was not distinguishable from zero — though the interval leaves room for a gain up to 4.67 percentage points — and cost per correct answer was $0.08665 with high effort versus $0.07662 with the term omitted. A contract census and raw-response probes further document that omitting the effort term means different things across models and even within a single provider, with claims explicitly bounded to the model, task, and collection date.
Cross-Model Memory Transfer via Target-Side Reader Adaptation
Engram-style hashed memory sits between retrieval-augmented generation and weight-baked adaptation: facts live in an external addressable table that a small learned reader consumes. The question studied here is what actually carries the value when such a table is detached from its source model and bolted onto a different backbone — the frozen memory or the reader that reads it. Ablations show both learned content and correct addressing matter, but the table is only useful through a reader aligned to the target model; a dual-layer, four-branch reader almost closes the gap between same-model and cross-model reuse, scoring 38.8 on average across downstream question answering, and a directly compatible provider reader gives substantial utility with no target-side training at all.
J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers
Fine-tuning turns a language model into a capable text classifier, but the decision criteria it learns stay locked inside the weights and only labels come out. J-Miner extracts them by aggregating vocabulary-aligned internal signals across layers and token positions into named text-level concepts, then fitting executable decision rules over those concepts using the classifier's own predictions as supervision. The rules reproduce up to 98.3% of source-classifier decisions and beat equally compact rules learned from raw input words by 6.0 to 29.5 percentage points of behavioral fidelity; distilling them into standalone students with about 1/24 the parameters retains 99.8% of the source classifiers' mean task accuracy.
Children, but not language models, show accelerating returns in word learning
Vocabulary growth in young children is usually modeled as steady evidence accumulation over time, but a reanalysis argues it is better described as accelerating accumulation, where each additional unit of linguistic experience teaches more than the one before. Language models — including ones trained on child-directed speech — show no such acceleration, instead exhibiting constant proportional returns on new data, consistent with scaling laws. The authors point to children's increasingly efficient use of their input as a candidate explanation for how they learn from many orders of magnitude less data than a language model needs.
A decodability criterion predicts when hidden-state selection beats majority voting in large language models
Combining several sampled answers from a large language model by majority voting breaks down on hard questions where the samples share correlated errors, and reading a correctness signal out of hidden states is an alternative whose reliability has been unpredictable. CASE (Correctness-Axis SElection) trains a linear gate on the answer-token hidden state to pick the highest-scoring candidate, paired with decodability, a leakage-free measure of how well that gate ranks correct candidates above incorrect ones for a given model and task. Decodability predicts the accuracy gain of selection over voting with a Pearson correlation of 0.75, and CASE beats voting by up to 19 points on medium-difficulty questions and 16.8 points on hard ones. The authors also show that a conventional probe only appears accurate because of question-identity leakage, which vanishes under question-grouped evaluation.
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn
Judging whether an assistant matches its explanations to a user's actual understanding is hard because existing user simulators do not model user knowledge or how it evolves during a conversation. KNOWSIM simulates users with an explicit knowledge state — a graph of information units connected by prerequisite relationships that updates under rules grounded in learning theory — and computes knowledge gain, delivery calibration, and cognitive overload from the resulting trajectory. Validated against 705 human-AI sessions across two domains stratified by knowledge level, its rankings agree with human judgments 73-74% of the time, ahead of three baseline simulators; applied to 9 LLMs it shows that which model is best depends on the user's knowledge level.
Polaris: Learning to Generate Table Descriptions from Retrieval Feedback
Table-centric tasks such as NL2SQL usually start with keyword search over LLM-written table descriptions that were optimized for fluency rather than for being retrieved. Polaris generates several candidate descriptions per table, ranks them by their actual BM25 retrieval effectiveness against existing query-table relevance judgments, and fine-tunes the generator on the resulting preference pairs with Direct Preference Optimization (DPO), after first expanding abbreviated table and column names to reduce vocabulary mismatch. It outperforms the state-of-the-art AutoDDG system, often by a significant margin, demonstrating that retrieval benchmarks can be repurposed as training supervision for retrieval-oriented metadata.
Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics
Ordering post-training data from easy to hard helps on some reasoning tasks and not others, and the analysis here traces that variation to the optimization dynamics each schedule induces. The authors formalize how training at one difficulty level affects performance at another as Relative Transfer, then derive Transfer-aware Dynamic Curriculum Sampling, which continuously reweights the sampling distribution according to estimated transfer during training. Across multiple reasoning benchmarks, model scales, and training paradigms the dynamic scheme consistently outperforms representative fixed scheduling strategies, and cross-difficulty transfer is offered as a unified, optimization-based account of when curriculum learning works at all.
Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking
Evaluating machine-generated scientific hypotheses usually relies on prompted large language model judges or semantic similarity, both of which can favor familiar ideas over novel ones. The alternative scores a hypothesis by the model's intrinsic confidence via a logit-based energy score, benchmarked across seven language models on 1,323 papers from 12 disciplines, each paper's true hypothesis competing against fifteen incorrect alternatives. Intrinsic scoring reached 33.0% Hit@1 pooled across scorers versus 16.6% for prompted listwise ranking, with the best configuration — a 1-billion-parameter model using energy scoring — reaching 53.1%, though that figure was the maximum over 14 model-by-scorer combinations selected after the fact.
What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?
Tokenizers are usually frozen before training even though the choice shapes how well a model handles different languages and scripts. Comparing tokenizer-free architectures — state-space language models (SSLMs) and H-Nets — against fixed subword tokenizers across 18 typologically and script-diverse languages shows joint optimization produces fundamentally different vocabularies: SSLMs recover morphologically aligned, contextually efficient units, while H-Nets favor byte-level efficiency with long tokens that barely overlap standard subword vocabularies. Agglutinative languages show the most dynamic segmentation during learning, and pretrained-then-finetuned BERT models using SSLM pretokenization lower perplexity while staying competitive downstream.
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration
Long-context prefill is dominated by the quadratic query-key score computation, and existing accelerations either push the whole attention through one low precision or drop token interactions entirely. TileMix instead makes precision a spatial routing decision: the score matrix is split into hardware-aligned tiles, routing bits are packed into compact bitmasks, and each tile group runs in FP16 or INT8 while both paths update one shared online-softmax state, so dense token connectivity survives and no training is needed. It supports grouped-query attention, variable-length batches, and INT8 key/value caches, and on LongEval, LV-Eval, and A100 prefill benchmarks with LLaMA, Qwen, and Vicuna it recovers the long-context quality lost to uniform INT8 while beating FP16 prefill throughput.
PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
Writing GPU kernels that exploit architecture-specific PTX instructions is a plausible target for code-generating models, but there has been no way to measure whether they actually do it. PTXBench scores functional correctness, whether the intended target instructions really execute at runtime, and speedup against frontier libraries on general matrix multiply and attention workloads across H100 and B200 hardware. Results are uneven — success rates drop sharply on complex attention backward passes, emitting the target instruction often fails to yield competitive performance, and no evaluated model consistently matches frontier libraries; supervised fine-tuning of Qwen3.6-27B with repair-conditioned training helps on several tasks but generalizes inconsistently, with data coverage, balance, and reasoning-teacher quality mattering as much as dataset size.
An Investigation of Translationese in the Generations of Multilingual Large Language Models
Text rendered from another language carries measurable traces of that process, known as translationese, and it is unclear whether multilingual large language models (MLLMs) leave the same signature when generating directly, as if translating internally from English. Generations in five languages are scored with established translationese indicators against both non-translated and human-written baselines, using high-accuracy classifiers, analyses of variance over individual linguistic features, and human annotations collected for German and Spanish. The design isolates how much translationese appears in MLLM output and which linguistic features separate it from the interference produced by direct translation, rather than conflating it with other sources of cross-lingual transfer.
When to Review: Spaced Repetition for Continual Pre-Training of Language Models
Continual pre-training has to absorb new data without erasing old knowledge, but replay schemes typically fix a global old/new mixture and sample it uniformly, ignoring that examples differ sharply in how fast they are forgotten. Spaced Repetition Training (SRT) reframes replay as adaptive review scheduling, keeping per-example review state and using the SM-2 (SuperMemo-2) spaced-repetition algorithm — with per-example perplexity mapped to a recall-quality signal — to decide which historical examples return and when, while leaving the model, objective, and optimizer unchanged. On temporally separated Wikipedia and code corpora it recovered 5 to 37 percentage points of the old-knowledge accuracy lost to naive continual pre-training across model scales without sacrificing new-knowledge acquisition, and at larger scale preserved broad benchmark performance that naive training and uniform replay degraded. Vision and tabular experiments suggest the scheduling principle transfers beyond language given a suitable recall signal.
MoNe: Modular Neural Memory for Efficient Long Context Inference
MoNe attaches a lightweight modular neural memory to any frozen pretrained Transformer to serve long contexts without retraining. It reads context in fixed-size segments through test-time learning of fast-weight memory networks with layer-localized gradient updates, and at inference generates keys and values from the query tokens alone, never re-reading context tokens, which decouples query cost from context length: O(N) preprocessing and O(1) queries with peak GPU memory that does not grow with N. At 128K tokens this cuts both compute and peak GPU memory by roughly 80% compared with in-context learning, at 6.4% parameter overhead, and the memory generalizes well beyond the backbone's native window on needle-in-a-haystack and word-extraction tasks from RULER, where in-context learning degrades sharply.
DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval
A single decoder-only LLM can both expand an underspecified query and encode text as dense vectors, but training it end to end creates a moving target: the retrieval gradients meant to improve expansion also shift the document embeddings that serve as retrieval targets. DEPT (Document Embedding Preservation Tuning) holds tuned document embeddings close to cached initial ones while letting retrieval gradients pass through straight-through decoding into the generator, converting joint query-document movement into query-side adaptation against approximately stable, whitened document embeddings that support index reuse and online hard-negative mining. With Qwen3-4B-Instruct-2507 and LLaMA-3.2-3B-Instruct on five BEIR datasets it improves average retrieval quality over training-free, independently trained, and staged unified baselines, with ablations isolating the contributions of preservation, whitening, end-to-end expansion training, and online negatives.
LLM-Derived Preference Judgments Are Not Self-Consistent
Agents that turn a person's natural-language preferences into numbers — asking a model how much someone would pay for a given flight — implicitly assume those judgments are self-consistent, meaning one utility function can reproduce them. The authors build statistical tests and interpretable measures of how far stated willingness-to-pay and indifference answers depart from the best-fitting consistent utility function, then apply them to flight, apartment, and hotel scenarios across six LLMs. All six models show large, persistent inconsistencies, implying that LLM-derived cardinal preference judgments cannot be faithfully summarized by a single utility function.
Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals
Most hallucination detectors operate at the answer or sentence level, which cannot localize which spans of a generation are fabricated or support fine-grained intervention. InnerExpert exploits signals available only in Mixture-of-Experts (MoE) models — router entropy, expert disagreement, and expert usage patterns — combining them with standard transformer features into compact per-token vectors classified by a lightweight detector trained on labels from an LLM-as-a-judge pipeline, so no manual annotation is needed. Across five datasets and two MoE architectures it outperforms existing methods, reaching 0.91 answer-level and 0.76 token-level AUROC from a single forward pass.
What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
Teams forced off deprecated commercial model versions typically decide using aggregate benchmark deltas, which compress heterogeneous per-item behaviour into one net figure. The authors query 900 public benchmark items covering graduate-level knowledge, olympiad mathematics, and instruction following 50 times each across three consecutive GPT-5.4-to-GPT-5.6 upgrades, classifying every item as reliably improved, reliably regressed, practically equivalent, or inconclusive under false-discovery-rate control and calibrated against a label-permutation null. Reliable improvements and regressions coexist in all nine migration-benchmark cells: an edge with a 7.3-percentage-point aggregate gain still contains up to 8.3% reliably regressed items, and edges with aggregate losses hide up to 10.7% improved ones. Strict versus loose instruction-following scoring also diverges, shrinking a 3.9-point regression to 0.04 points.
Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
Three frontier mixture-of-experts models with 3.6–4.0B active parameters were fine-tuned to reason in Greek, and accuracy benchmarks moved almost not at all — changing only the random seed shifted scores by 7.7 points, more than any data or recipe effect measured. The behavioural changes were large: base models produced Greek reasoning in 0 of 1,000 traces even for Greek questions, whereas every supervised fine-tuned (SFT) checkpoint reasoned in the question's language on roughly 98% of items, one family at three times fewer tokens, with better judged grammaticality and general ability within a few points of base. SFT could not repair its own defects — a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit instruction to think in English is obeyed under half the time — while pre-registered reinforcement learning with verifiable rewards fixed the first two outright (format fallback 24% to 2.5%, leakage 3.5% to 0.0%) against a flat random-reward control and moved the third by 9.1 points. The authors release five checkpoints plus six length-decorrelated behavioural metrics, and report six cases where their own instruments misled them.
Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility
Systems increasingly tailor retrieved evidence to the identity of the downstream model that will read it, which only pays off if reader differences form reusable structure rather than input-local interactions. Holding query, evidence, task, scoring, and intervention fixed across nine readers in a retrieval-augmented generation setup, readers disagree on the sign of an intervention's effect in 33% of jointly affected cells, and reader-by-query interaction explains 29.8% of utility variance against an 8.4% permutation null. Ordinal reader geometry — how a reader ranks evidence — stays stable across four independent settings, but signed help-or-harm direction is weak in open-ended question answering and strong only in binary fact-checking. Critically, stable ranking similarity fails to predict whether an intervention transfers between readers, so preference cannot be treated as a license to reuse help/harm decisions.
Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
When a user states a belief grounded in false information, language models often fail to acknowledge it, and this work shows the failure is largely a function of the verb used. Testing 10 models across 18 epistemic expressions reveals an accuracy gap between factual and false content ranging from +50% for "I vaguely remember" to -14% for "I seriously doubt". The cause is task confusion: models default to fact-checking the underlying claim instead of tracking the stated belief, chains of thought that explicitly fact-check score worse on false information, and a single instruction can reverse the failure across verb families. Attention analysis shows models attend more to false beliefs they fail to confirm, but suppressing that attention at decoding time recovers accuracy only partially and only in some models.
Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses
Evaluating language models with tests designed for humans implicitly assumes both are drawing on comparable underlying abilities. To check that assumption, responses from humans and six models on quantitative reasoning and chemistry assessments were run through separate Exploratory Factor Analysis, and subject-matter experts then blindly tried to attach pedagogical meaning to the resulting factor graphs. Experts interpreted most human-derived factors but could not ascribe meaning to any model-derived factor in quantitative reasoning, and made sense of only half of them in chemistry, suggesting models solve these assessments through statistically opaque mechanisms unlike human reasoning constructs.
From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector
Public institutions choosing a language model get little help from mainstream benchmarks, which are English-language, US-centric, and focused only on task accuracy. MÖVE is an evaluation framework aimed at the German public sector that adds three governance dimensions: estimated energy consumption, provider transparency, and knowledge of German political party positions. No model leads on all three — estimated energy consumption varies more than 60-fold across models and is not explained by size alone — information disclosure differs systematically by provider, and European models show no advantage on German party positions, so procurement cannot rely on performance rankings alone.
Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints
Probing studies routinely report that models encode some structure internally, but rarely test whether that encoded information actually drives behavior. Using parametric computer-aided design constraints as a controlled geometric testbed, six frozen decoder-only models are audited on four properties: linear decodability, forced-choice generation, activation-level influence, and behavioral steerability. Pretraining clearly improves decoding of local pairwise geometric relations, but sketch-level degree-of-freedom status is already highly decodable from randomly initialized weights and barely improves with pretraining; more strikingly, decodable information often fails to surface in generation, activation-restoration effects vanish while decodability persists across depth, and mean-difference steering does not reliably control outputs.
BayesPrompt: human readable prompts that make sense
Prompt optimization methods that minimize the perplexity of a target answer tend to produce pseudoprompts — unintelligible token strings that work but cannot be read or reasoned about by humans. The authors argue this stems from the ill-posedness of the optimization problem itself and reframe prompt recovery as Bayesian posterior inference over prompts, giving an efficient sampling algorithm. BayesPrompt produces prompts that are both low-perplexity and human-readable, with reported improvements over state-of-the-art alternatives across a range of metrics on a real dataset.
Grading Needs a Rubric, Not Intelligence
Tests whether grading open-ended exam answers needs an expensive judge model or merely an explicit rubric. In the any-to-bench design, a frontier model reads source documents once at ingestion to extract each question and its rubric, after which six cost-efficient configurations from two model families at three reasoning-effort levels both answer 24 questions and grade every answer sheet three times, producing 3,456 per-question grades. Answer identity explains 95.6% of score variance while judge identity explains only 0.2%, and raising a judge's reasoning effort moves assigned scores by at most 0.006 of full marks versus 0.143 for a writer's effort. Ablations locate the effect in the official answer rather than the criteria and levels: stripping criteria changes nothing measurable, while removing the official answer too drops the intraclass correlation from 0.888 to 0.628, inflates scores, and makes judge effort matter again.
Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds
Sidesteps attention-weight interpretability, which routing artifacts such as attention sinks distort, by analyzing the geometry of hidden states directly: continuous similarity matrices over long-context representations are sparsified into unweighted graphs whose connectivity between disjoint semantic anchors is then traced. Across two architectures a sharp topological transition appears — early syntactic layers stay entirely fractured, while deep reasoning layers compress large conceptual distances into navigable paths of at most six semantic hops, a small-world structure. Applied to zero-shot hallucination detection for retrieval-augmented generation (RAG) on the RAGognize dataset, factually grounded generations stay roughly three hops from their source context whereas hallucinations induce topological collapse, yielding a purely geometric reliability signal.
Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
Reference-free LLM judging on objective tasks forces a choice between fast parametric verdicts that may hallucinate and tool-augmented verdicts that cost more and need a policy for when evidence is used, with neither offering formal control over the error rate of accepted verdicts. The proposed framework calibrates uncertainty thresholds on a held-out set with finite-sample Clopper-Pearson intervals so that the false discovery rate among accepted verdicts stays below a user-specified level with high probability; instances the parametric mode is not confident enough about are routed to a retrieval-augmented mode that gathers web evidence and re-evaluates under a second calibrated threshold, and the guarantee carries over to this two-threshold routing without extra assumptions. Across open-domain question-answering benchmarks and judges of varying scale, it holds the target error rate while accepting substantially more instances than single-mode baselines.
Chain-of-Experience for Continual LLM Improvement
Conventional evaluations measure zero-shot output and ignore whether a model improves through interaction at inference time. Chain-of-Experience names the setting in which a model accumulates experiential traces over iterations using self-feedback or environmental signals such as correctness labels and public coding-test pass rates, evaluated across math, coding, and knowledge tasks with eight models including GPT-5, Gemini-2.5 Pro, and Claude-4.5 Sonnet. Iterating on experience consistently beats feedback-free baselines, yielding a 5.6% overall improvement alongside 19% lower API cost, with combined feedback channels adding further gains and better accuracy per token than other test-time strategies. Stronger base models improve more, behavior stays robust under weak or spurious feedback, and most of the benefit arrives in the early iterations.
TokEval: A Tokenizer Evaluation Suite
Tokenizers are usually picked with little more scrutiny than fertility and compression rate, partly because it is unclear which tokenizer properties drive which downstream abilities. TokEval adds linguistically and structurally grounded metrics — UTF-8 character boundary integrity, digit place-value boundary alignment, line-break handling — and validates them through controlled pretraining runs that vary only the tokenizer's training mixture, pretokenization strategy, and training algorithm, evaluated on bits-per-byte plus language, math, and code benchmarks. Information-theoretic metrics predict language modeling quality at Spearman correlations up to 0.80, while structure-sensitive metrics track task accuracy instead, suggesting intrinsic measurement can replace some expensive pretraining sweeps.
6 more specialized papers
- Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs Florian Braun
- Structuring Semantic Embeddings for Principle Evaluation: A Prototype-Guided Contrastive Learning Approach Che Shen, Junwei Su, Lingpeng Kong et al.
- L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages Rinit Jain, Tirthraj Mahajan, Advait Joshi et al.
- A Pre-Specified Construction-Confirmation Test of Operation-Level Causal Transfer Across Finite Isomorphic Symbolic Domains Xinyi Shan
- Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation Seung-Won Seo, Won Ik Cho, Yongmin Yoo
- Q-Interference: Memory-Efficient Phase-Aware Quantum-Inspired Attention Emama Nahid, Tahmid Imtiaz Imu, Huayue Gu et al.
Theory 74
Quantifying Depth Sufficiency in Residual Neural Networks: A First-Order Criterion
Under a fixed function-preserving residual-growth protocol — specified insertion locations, residual families, zero-output initialization, and zero-state first-order updates — the question of whether a trained network is already deep enough gets a formal criterion. First-order residual depth saturation is defined as the absence of any strictly improving admissible insertion, and the paper proves residual non-degeneracy is necessary and sufficient: additional depth has first-order value exactly when conditional activation gradients project nonzero onto at least one admissible residual tangent space, a boundary invariant under regular local reparameterizations. Empirically, across ResNets, GPT-2-style models, and continued-pretrained Pythia checkpoints, the maximum activation-gradient norm falls toward a low-signal regime as depth grows, and grown models match training from scratch, supporting gradient magnitude as a conservative depth diagnostic.
When Does the Best Sampling Temperature Rise with the Budget? Sufficient Conditions for Pass@k
The familiar pattern where the best sampling temperature for pass@k is low at small budgets and higher at large ones is not a property of the metric itself — for one fixed task the maximizing temperature is independent of k — so it must arise from aggregation across tasks of varying difficulty. Defining a conditional log-success response that measures how a task's success probability reacts to temperature given its current success rate, the paper proves that if that response is nonincreasing in current success probability, the normalized temperature derivative of aggregate pass@k is nondecreasing in k, making derivative signs nested across budgets and single-peaked curves' maximizers nondecreasing in k. The mechanism is identified as a monotone-likelihood-ratio power tilt toward lower-success tasks, developed into a closed-form two-stratum phase diagram and an exact Beta(2,k) kernel that concentrates at one-sample success of order 1/k. No model was trained or queried; the contribution is a conditional theory whose assumptions remain to be tested.
A Reproducibility Study of Partial Residual Ablations in Pre-LN Transformers
The attention and feed-forward residual pathways in Pre-LayerNorm GPT-style transformers are ablated separately, together, and not at all, at 10M and 124M parameters. Removing the attention residual consistently collapses performance to the no-residual floor, while removing the feed-forward residual shows a reproducible recovery effect at 10M under a controlled 8-seed deterministic study; the 124M case remains unresolved because of substantial seed variance. The write-up documents a measurement confound in runtime gain scaling that was identified and corrected mid-investigation, reports the failed intermediate reproduction alongside the successful replication, and proposes a cross-position routing hypothesis for the asymmetry while separating confirmed findings from open questions. Code, configs, checkpoints, logs, and non-reproducing runs are all released.
CFR without Unbiasedness: Deterministic Guarantees for Persistent Public-Chance Schedules
At a finite public-chance cut, counterfactual regret minimization (CFR) must decide how many chance outcomes to evaluate before each regret update; persistent partial evaluation walks a fixed without-replacement order so every outcome is covered once per epoch, but its feedback is conditionally biased because earlier batches shape the strategy profiles later batches see. A deterministic target-transfer theorem bounds full-cut exploitability by the regret on delivered feedback plus a public-debit term coupling prefix coverage discrepancy with motion along the realized strategy path, which yields convergence for signed regret matching and RM+ under consecutively balanced schedules and converts an execution trace into a numerical exploitability certificate. On two heads-up no-limit hold'em turn endgames, persistent ordering substantially outperforms fresh reshuffling despite identical per-epoch coverage, and partial coverage wins every shallow matched-budget comparison until a crossover between 32 and 64 full-cut outcome budgets, after which complete coverage dominates.
Is Grokking a Loss of Normal Hyperbolicity of the Interpolation Manifold?
Recent work frames the post-memorization phase of grokking as a fast-slow dynamical system in which weight decay drifts the network along a zero-loss interpolation manifold, raising the question of whether the sharp generalization transition is a bifurcation where a normal restoring direction goes flat. The proposed diagnostic is the smallest nonzero singular value of the residual Jacobian, which under squared loss equals the slowest normal restoring rate of the manifold. On a two-layer ReLU network grokking modular addition, that singular value does not collapse at the transition but instead attains its largest values during it, holding across five seeds with no subspace-local collapse in the six smallest singular values either. The authors present this as preliminary evidence for smooth contraction rather than a bifurcation, while noting a single gradual-transition setting under Adam cannot rule one out.
Optimal Watermark Localization in Mixed-Source Large Language Model Texts
Text that mixes human and model writing carries watermark evidence at only some token positions after rewriting, insertion, deletion or paraphrasing, and prior work only asks whether a watermark is present globally rather than where it survives. Localization is cast as token-level multiple testing over pivotal statistics with a latent per-position indicator, and an asymptotic analysis indexed by signal sparsity, next-token concentration and effective-vocabulary growth yields a sharp detection boundary plus phase transitions for discovery and classification. Discovery is provably strictly harder than detection, and consistent classification is impossible anywhere in the parameter regime for coordinatewise pivot-based rules. An adaptive thresholding procedure that estimates the surviving watermark fraction from data, without knowing the exponents or the time-varying next-token distributions, attains the optimal discovery boundary and near-optimal power, with simulations and model-generated text under common edits supporting the theory.
Learning reshapes power-law anisotropy in internal representations
Power-law anisotropy in internal representations shows up everywhere from large language models to mouse cortex, but how it arises from input structure and training has been unclear. The authors exactly solve the learning dynamics of a wide two-layer linear network in a teacher-student setup with power-law input and teacher spectra, and find that in the feature-learning regime the local power-law exponent evolves nonmonotonically during training, passing through up to four distinct asymptotic regimes across modes and time, while in the lazy regime it barely moves at all. Numerical experiments indicate similar exponent dynamics in more realistic nonlinear networks, pointing to interaction between input statistics and task structure as the general mechanism.
Towards a theory of inference-time alignment with unknown rewards
Inference-time alignment is usually analyzed assuming access to a decent reward estimate; this work instead treats it as weak-to-strong learning where everything, including the reward, is learned from data, and a reasonably good reference policy must be turned into one that emits a good response with arbitrarily high probability. Following the PAC (probably approximately correct) learning template and allowing multiple acceptable responses per prompt, the authors define a new combinatorial quantity called the alignment dimension and show a reward class is alignment learnable exactly when that dimension is finite. The learning procedure runs the ordinary one-inclusion graph algorithm as a tournament over all pairs of label sets where neither contains the other.
Guaranteed Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group
When a multimodal system can pay to acquire additional inputs at inference time, the acquisition policy itself decides which input pattern a given case ends up in, which breaks the usual assumption behind conditional calibration that the grouping map is fixed independently of the calibration data. This work characterizes when pattern-conditional guarantees survive that dependence and gives two finite-sample constructions — threshold-free routing calibrated at the terminal pattern, and simultaneous certification of whole policy-pattern pairs so calibration data may also select the deployed policy — along with a counterexample showing a guarantee proved for a calibration-independent grouping need not carry over. On a staged electrocardiogram lead protocol the certified RouteCert policy answers 71.2% of held-out patients at 7.4% disagreement with the cardiologist for 48.8% of the full acquisition cost, and on masked multimodal benchmarks per-pattern certification holds worst-pattern selective risk at 0.034 where a pooled design reaches 0.145 against a 0.10 cap at comparable coverage.
Sparse Prototype Code Underlies Classification and Prediction Across Modalities
High-dimensional neural representations are hard to interpret, but classification tasks appear to induce a shared geometry across vision, audio, and language models. The analysis shows that within-class variability is not isotropic noise: its classifier-relevant part correlates with the class centroid and with the centroids of competing classes, which supports an analytical mean-field theory including a renormalized class radius that corrects for non-Gaussian statistics. The theory predicts classification accuracy across architectures and modalities using only a small set of centroid coordinates for the true class and its strongest rivals, and the relevant geometric quantities improve with model scale. The sparsity of the required description connects the framework to sparse-feature extraction methods such as sparse autoencoders.
Generalised Transportability via Causal Abstractions
Classical transportability theory decides whether a single causal query can be carried from a source population to a target one, but it answers one query at a time, returns a formula rather than a value, and says nothing when the query is not transportable or when no target data exist. Recasting the problem through Causal Abstraction theory, the work asks whether one map aligns source and target across all interventional behaviour, and characterises when such a map exists in Markovian and semi-Markovian settings; when it does, every target query transports simultaneously. The main contribution covers the approximate case: the best available map turns abstraction error into certified intervals for queries, formulated as distributionally robust optimisation over mechanism and environment perturbations. On synthetic benchmarks and a real ecological dataset, the certified intervals bracket the true interventional quantity.
On Stopping Rules and Spatial Adaptation for CART
Regression trees fit by CART pair a greedy splitting rule with a stopping rule, and while splitting has been studied closely, the statistical role of stopping has not. Under spatially heterogeneous and anisotropic smoothness plus structural assumptions on the regression function and covariate distribution, CART with a minimum impurity decrease (MID) stopping rule and a suitable threshold is shown to attain pointwise rates that are minimax up to logarithmic factors, holding simultaneously at all points in the domain. The widely used minimum leaf size stopping rule provably cannot achieve this spatial adaptation, giving a concrete statistical reason to prefer MID and a theoretical account of why CART works well in practice.
Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception
Language models are ranked by held-out per-token cross-entropy risk, the same quantity scaling laws are fitted to, and the argument here is that this risk cannot be consistently estimated by any estimator, the holdout average included. The obstacle is topological: across the space of possible worlds — a data-generating distribution paired with whatever model gets trained — states of finite and infinite risk lie arbitrarily close to one another, and whether a model's risk is estimable is a tail property of the distribution its weights induce that no finite sample reveals. Inconsistency survives bounding expected sequence length and restricting to full-support models, and in that restricted setting the bad states are dense. Two escapes are offered: flooring next-token probabilities under a bounded context window, which gives a statistical rather than merely computational rationale for finite contexts, or reporting risk only below a threshold fixed in advance, which restores consistency at the cost of redefining the estimand.
Second-Moment Memory in Coordinatewise Adam
The moving average of past squared gradients in Adam's denominator carries an optimization cost that has not been quantified. Analyzing a simple two-point oracle, the expected positive normalized update decays as $O(M_2^{-1/2})$ after an initialization transient, where the memory length $M_2$ equals $(1-\beta_2)^{-1}$, and this directional bound converts under the stated memory and stepsize scaling into an average-stationarity lower bound of the same order on a smooth convex problem. Long second-moment memory can therefore slow progress toward the optimum even when gradient noise has finite variance, rather than only under heavy-tailed noise.
How Many Samples Are Needed to Determine Causal Direction? Sharp Minimax Bounds for Bivariate LiNGAM
Classical LiNGAM theory shows that independent non-Gaussian disturbances make the causal direction between two linearly related variables identifiable in the population, but says nothing about how many samples are needed when the effect is weak or the disturbances are nearly Gaussian. The result here is a sharp local minimax sample complexity that scales as $\log(1/\delta)$ divided by $d_\beta^2 + \beta^2\nu^2$, where $\beta$ lower-bounds the structural coefficient, $\nu$ measures distance from Gaussianity, and $d_\beta$ captures the residual scale uncertainty, which also separates the regimes where identification comes from non-Gaussian dependence versus from covariance alone. The author notes the proof was generated in a two-hour session with GPT-5.6 Sol in Codex's Ultra mode, with the human supplying the prompt and checking, revising and polishing the result.
Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive
Continual-learning regularizers such as EWC limit forgetting by penalizing parameter drift weighted by diagonal Fisher importance, which looks flexible but summarizes each layer's curvature poorly, discarding the top-eigenvalue information that actually governs forgetting. Assuming a block-diagonal Hessian — the layer-level counterpart of EWC's own diagonal assumption — the authors prove that forgetting decomposes into per-layer terms weighted by each layer's top Hessian eigenvalue, that diagonal Fisher weights cannot recover that eigenvalue (two layers with identical Fisher averages can have top eigenvalues differing by a factor as large as the layer width), and that uniform regularization sacrifices new-task performance in proportion to the layer condition number. The practical recipe that follows is to protect early layers strongly and let deeper layers move, which improves average accuracy and forgetting metrics for both EWC and SLCA.
Information Geometry of Message Passing
Starting from the Bethe free energy and constraining one edge marginal of a Forney-style factor graph to an exponential family, the natural-gradient stationary condition of variational inference turns out to have a purely edge-local form: at a stationary point the edge's natural parameter equals the sum of two projected messages, one from each incident factor, where each projection is the natural-gradient projection of the exact belief-propagation log-message at the receiving marginal. The resulting natural-gradient message passing (NGMP) lets every edge carry its own exponential family and makes a factor's outgoing message depend on which marginal receives it. Unlike variational message passing, it keeps whatever part of the exact message the receiving family can represent instead of averaging the factor under neighboring beliefs; the two agree when edge uncertainty entering a non-conjugate factor vanishes, and experiments on Poisson smoothing, heteroskedastic regression, and hourly ETTh forecasting show the advantage shows up mainly as better uncertainty calibration when that uncertainty persists.
Toward Optimal Second-Order Path-Length Guarantee for Adversarial Multi-Armed Bandits
Adversarial multi-armed bandit algorithms can adapt their regret to how much the loss sequence moves over time, and Bubeck et al. (2019) left open whether regret scaling with the second-order path length is reachable under bandit feedback. A more careful analysis shows that their existing algorithm, unchanged, already achieves regret of order K·log(KT) + √(K·log(KT)·(1+Q∞,2)), matching the known Ω(√(K·Q∞,2)) lower bound up to logarithmic and additive terms. An adaptive restart scheme with a path-length estimator whose increments are uniformly bounded removes the need to know the second-order path length in advance.
Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics
Autoregressive neural emulators of chaotic systems drift and blow up over long rollouts, and fixes for this have been largely ad hoc. Analyzing the Jacobian of the learned one-step update with respect to the state shows that inference-time error growth is governed by its spectral radius: direct-step architectures that map state to next state generically pick up eigenvalues above magnitude one, while integration-constrained models that predict a time derivative and integrate it with a higher-order scheme collapse the eigenspectrum onto the unit circle, giving neutral stability and a universal linear error-scaling law. The largest Jacobian eigenvalue thus serves as a cheap, architecture-agnostic predictor of short-term skill, long-term stability, and spectral bias without running a rollout, and a stability-promoting loss that regularizes this amplification improves both accuracy and robustness. The framework is demonstrated on 29 models spanning two architectures, several explicit and implicit integrators, and multiple losses on the Kuramoto-Sivashinsky system.
The Trade-off Between Covariate Dependence and Latent Structure in Representation Learning
Disentangled representation learning wants each latent dimension to track a distinct covariate, but unsupervised methods only enforce independence without semantic grounding, and supervised methods cannot pin down one-to-one latent-covariate alignment and latent independence at once when covariates are correlated. A unified supervised framework couples latent-covariate dependence with explicit constraints on latent structure, and the authors prove that enforcing either latent independence or exclusive one-to-one dependence costs alignment, yielding an ordered hierarchy of disentanglement regimes each with a closed-form latent-space transformation. The transformations are applied post-hoc to realign pretrained CLIP, DINOv2, and ViT representations and folded into informed factor analysis, giving controllable structured latents on simulated and real multi-omics data.
Probabilistic Circuits as Reasoning Machines in Artificial Intelligence (Part I)
A habilitation thesis synthesizing a decade of work on probabilistic circuits (PCs), a model class whose structural constraints make otherwise NP-hard probabilistic inference tractable. It argues for probability as a core language for AI — via its links to logic, information theory, human cognition, and optimal decision making — then shows how PCs compute marginals, conditionals, most probable explanations, and expectations exactly in polynomial time. Covered contributions span foundational theory, Bayesian learning of PCs, scalable implementations integrated with deep learning, hybrid models pairing PCs with intractable ones, and connections to symbolic machine learning; the cumulative second part is omitted.
On the Principles Behind Neural Network Optimizers
Adam remains the default optimizer for training large language models despite unsettled theory about when it converges and why it outperforms stochastic gradient descent (SGD) on Transformers. This thesis identifies a problem-dependent phase transition — with properly chosen, batch-size-dependent hyperparameters Adam converges, while small second-moment-decay regimes can diverge — and traces its advantage on Transformers to the Hessian evolving toward a near-block-diagonal form with strong heterogeneity between blocks, which is proven to make a diagonal preconditioner effective; random matrix theory links that structure to consecutive multiplications of large matrix variables. The analysis motivates Adam-mini, which reduces Adam's memory footprint by 50% while preserving performance, and also offers a lens on newer optimizers such as Muon.
Spectral Gaps of Hit-and-Run and Coordinate Hit-and-Run
Sampling uniformly from a convex body is a basic primitive, and it had remained open to tie Hit-and-Run's convergence to Poincaré and Kannan-Lovász-Simonovits (KLS) constants the way Kannan, Lovász and Simonovits did for the Ball walk. Bounding the chain's spectral gap directly through functional isoperimetric constants — rewriting it via dual certificates, which surfaces the Babuška-Aziz constant from PDE analysis, rather than through conductance — gives a gap of order one over dimension-squared times the Poincaré constant. For nearly isotropic bodies this yields a mixing time of order n²·log n with only logarithmic dependence on the initial distance, improving the dimension dependence from cubic to nearly quadratic, and the same technique gives Coordinate Hit-and-Run a much improved n³ bound.
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve
The best known bounds on the matrix multiplication exponent come from combination loss analysis, a refinement of the laser method whose bottleneck is a hard nonconvex optimization problem. The problem is reformulated so it can be solved in a larger setting than before, attacked with a new optimization algorithm built on recent machine learning techniques, and then refined using AlphaEvolve. The result is an upper bound of ω < 2.371177, improving the previous best of 2.371339.
There is No Theoretical Curse of Multilinguality For Embedding Space Structure
The curse of multilinguality — the observed decay in per-language quality as a single model covers more languages — is usually treated as an inherent capacity limit of shared embedding spaces. Formalizing "perfect multilinguality" as two explicit conditions on monolingual quality and cross-lingual alignment, the authors prove that the minimum embedding dimensionality needed to satisfy both grows only logarithmically in the number of languages. That implies the empirical curse comes from real-world data and training conditions rather than from geometry, a conclusion they support with a small-scale empirical study.
Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making
How cautious an agent should be while still learning its environment is answered by tying caution to epistemic uncertainty: RATTL (Risk-Adversarial Total-Reward Learning) holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior, so behavior interpolates continuously between worst-case robustness and risk-neutral reward maximization as evidence accumulates. The construction follows the duality behind Entropic Value-at-Risk, which converts a choice of risk level into a choice of ambiguity radius. Beyond well-posedness under transience and compactness, the main result is a 'Safety Sandwich': the RATTL value lies between the uninformed robust value and the full-knowledge optimum, with the gap vanishing as the posterior concentrates; in a canonical binary-hazard instance the criterion reduces to Conditional Value-at-Risk at a level set by posterior entropy, and the agent defers the efficient action until a sharp identification threshold.
Evaluating the Diversity of AI-Generated Content with Diversity Profiles
Generative model diversity is usually reported as one scalar — embed the samples, compute pairwise distances, aggregate — but different scalars encode different inductive biases and can rank the same sample sets in contradictory orders. The argument that the measurement is intrinsically under-specified rests on an axiomatic result that no representative scalar metric satisfies all the desirable properties simultaneously, plus evidence that high-dimensional representation spaces produce concentrated, modality-dependent distance distributions. In place of scalars the paper proposes diversity profiles: curve-valued, condition-aware summaries that sweep a parameterized diversity family across thresholds, scales, exponents, or orders, making visible whether a comparison holds across resolutions or depends on an arbitrary parameter choice.
A Theoretical Framework for Parallel Lifelong MAPF Using Group Decentralized Planning
In lifelong multi-agent path finding, agents continuously receive new destinations and must route without colliding; the leading Rolling-Horizon Collision Resolution approach produces high-quality plans but costs too much compute to scale past modest agent counts. Borrowing tools from Locally Interdependent Multi-Agent Markov decision processes, the authors first prove that the rolling-horizon method is near-optimal in a discounted formulation, then extend it to GD-RHCR, which partitions agents by a transitive communication scheme and plans each partition in parallel. Both variants achieve exponentially close-to-optimal guarantees, establishing a duality between time-based horizon truncation and space-based partitioning, and the partitioned version sustains high throughput at larger agent counts with substantially lower per-plan cost.
An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models
Examines what acceptance actually certifies in the Code World Model setup, where an LLM synthesizes an executable world model, a classical planner searches it, and the model is accepted for reproducing sampled transitions. The pipeline's danger factors exactly: N independent gate rollouts all miss a critical event of probability r with probability (1-r)^N, so acceptance certifies sample consistency and no more. On three hybrid continuous-control instruments the accepted mode-blind model is exploited by the planner at a regret of nearly the entire attainable return, and a proved Lipschitz localization budget explains why continuous discrepancies must occupy detectable volume while discontinuous reset modes pay no such price. With real synthesis, GPT-5.x repairs an omitted one-dimensional clamp in 105 of 111 mode-containing draws but recovers the two-dimensional region rule in 0 of 156, and re-scoring all 1,034 artifacts finds the gate provably informative on only about two percent of the exploited planner's queries.
Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System
Symbolic music tokenizers often imitate language modeling by treating chords, motifs, and phrases as reusable units, but the argument here is that tokenization's advantage comes from compression, which requires coordinates in which recurring regularities form stable, predictable conditional distributions rather than simply larger combinations. The Effectiveness-Losslessness Framework defines two boundaries: decoupling and denesting build coordinate interfaces that expose predictive regularities, while tokenization must stop before context-dependent relations are fixed, leaving those to model states. Controlled symbolic-music experiments find that sequence compaction alone does not guarantee predictive compression, whereas preserving contextual freedom lets higher-order musical organization emerge without explicit structural labels — architectures transfer across modalities, tokenization interfaces do not.
44 more specialized papers
- PathFinder: Joint Decompositions of Linked Multimodal Datasets Ying-Qiu Zheng, Alex Fung, Stephen M Smith et al.
- A Deep Learning Model for Spatially Clustered Data via Differentiable Cluster Assignment Kexuan Li, Weidong Ma
- Spinning Conformal Correlators from Neural Networks Manas Dogra, James Halverson, Joydeep Naskar
- Lipschitz Bandits with Arbitrary Feedback Delays Yuhao Liu, Yu Chen, Longbo Huang
- Online Convex Optimization with Dueling Feedback Yiyang Lu, Hareshkumar Jadav, Mohammad Pedramfar et al.
- Sufficient Dimesion Reduction via Generalized Stein's Lemma Ye Tian
- Beyond Effective Sample Size: Effective Number of Proposals for Adaptive Importance Sampling Ali Mousavi, Victor Elvira
- The Physical Cutoff Does Not Restore Homogenization: Phase-Dependent Burning in the Strain G-Equation Michele Caprio
- Prediction Inference of Time Series with Standard ReLU Deep Neural Networks Kejin Wu
- Does 1/2-Tsallis-INF Also Work Well for Best-Arm Identification? Jingxin Zhan, Yuze Han, Zhihua Zhang
- High-Dimensional Nonparametric Change-Point Detection via Low-Rank Degree-Three Density Projection Guoqing Zhang, Zhaixin Chen
- Optimal Lower Bounds for Networked Information Aggregation Ambar Pal
- A Counterexample to the Tang Zhang Schatten Norm Conjecture and Sharp Positive Results Zijian Zeng, Houde Liu, Kurunathan Ratnavelu
- Inferential Evaluation of Surrogate-Derived Models under Covariate Shift Longtian Shi, Molei Liu, Doudou Zhou
- Generalized Linear Bandits with Memory Heesang Ann, Hyunjun Choi, Taehyun Hwang et al.
- $S^3$: A Smooth Simulation Surrogate for Optimizing Discrete Abstractions of Dynamical Systems Jordan Peper, James Mathias Gast, Vignesh Nanduri et al.
- A Banach-Space Theory of Markovian Halpern Iteration for Non-Expansive Maps Ege C. Kaya, Arda Fazla, M. Berk Sahin et al.
- Fiber Fingerprints of Hidden Learning-State Dynamics Qinyou Wang
- Operator-Theoretic Generalization Bounds for Multitask Deep Learning Mahdi Mohammadigohari, Thomas Borsani, Giuseppe Di Fatta
- EMS Coreset: An Efficient Expectation-Maximization Algorithm for Sinkhorn Coreset Haoyun Yin, Chuanhui Liu, Xiao Wang
- Coded Hankel Polynomial Chaos: Spectral Identification of Dominant Polynomial-Chaos Modes Zhiliang Deng, Xiaomei Yang
- Demystifying Oversmoothing in Sheaf Neural Networks: An Index-Theoretic Criterion Junwen Dong, Yuhan Peng, Hao Li et al.
- Beyond Peak Backlog: Conditional Energy and Temporal Geometry in Capacity-Constrained Delayed Bandit Optimization Anling Xiang, Yuwen Yang, Yang Shen
- Correlation Clustering with Random Partial Information Rajath Rao K. N., Jens Schl\"oter, Sami Davies et al.
- LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models Tom Splittgerber, Niklas Koenen, Marvin N. Wright et al.
- Coverage-Maximizing Multinomial Subset Routing under Operational Constraints Quan Zhou, Yiyan Huang
- Reference-free logged energy-oracle recovery for neural approximations of symmetric coercive variational problems: conforming Riesz reconstruction and archive-level selection Karim Bounja, Lahcen Laayouni, Boujemaa Achchab et al.
- Improved Regret Analysis for Parallel Gaussian Process Bandit Optimization Shion Takeno, Shogo Iwazaki
- Density-Reweighted Entropic Optimal Transport: Decoupling Geometry from Sampling Density Keyi Li, Yuval Kluger, Boris Landa
- Random Quadratic Form with random forcing: Metastable synchronization by noise Anna Shalova
- Learning to Price with Persuasion Maria-Florina Balcan, Tejas Pagare, Karan Singh
- The canonical facets of multi-separator polytopes Bjoern Andres, Silvia Di Gregorio, Jannik Irmai et al.
- Memory Is Communication: The Frontier Between Remembering and Signaling Yashar Talebirad, Eden Redman, Ali Parsaee et al.
- Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions Xiao Wang, Tomohiro Hashizume, Pia Siegl et al.
- Expected free energy as an information constraint on the Bethe Lagrangian Wouter M. Kouw
- Maximum Tsallis Entropy Distributions for Robust and Efficient Sparse Learning from Correlated Data Kai Yang, Masoud Asgharian, Celia M. T. Greenwood
- Nonadaptive Learning in Robust Nonlinear Output Regulation Shimin Wang, Martin Guay, Richard D. Braatz
- Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression Tao Jiang, Minbo Gao, Shaowei Cai
- Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models Satpreet Makhija
- Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho et al.
- The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting Nazl{\i} Nur Karabulut, tanya Braun
- Adaptive Policy Portfolios for Robust Markov Decision Processes Kasper Engelen, Sebastian Junges, Guillermo A. P\'{e}rez et al.
- Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints Chainarong Amornbunchornvej
- Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation Hollis Robbins (University of Utah)
Other 55
p-Spin Glass Network Efficient Single-Batch Continual Learning
Sequence models typically need large memory footprints and large-batch stochastic optimization, which limits sample efficiency and continual learning. The proposed p-Spin Glass Network uses native ternary quantization to compress internal parameters 8x and exact implicit gradients to bound activation memory, and is claimed to match a Transformer baseline's asymptotic performance while using 8x fewer training sequences. The headline claim is smooth, monotonic convergence at a stochastic micro-batch size of 1, holding across both discrete subword tokens and long-horizon raw byte streams, which the authors present as removing the large-batch requirement for stable deep learning and a foundation for edge deployment.
What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models
Music foundation models are routinely frozen and used as feature extractors, yet which layer to tap is chosen by heuristic or papered over with multi-layer fusion. A layer-wise analysis of 12 models spanning masked, autoregressive, and contrastive pre-training correlates label-free geometric and transformation-based representation-quality metrics with downstream performance on 15 tasks. Several metrics track layer quality for genre classification, emotion recognition, tagging, and beat tracking, but all of them fail on tonal tasks such as key estimation and chord recognition, prompting a new pitch-transposition equivariance measure that gives a consistent tonal-quality signal across model families. Used for layer selection, these intrinsic metrics match or beat trainable multi-layer fusion, especially when downstream data is limited.
Disentangling Homophily and Rarity: Explaining Failure in Graph Neural Networks
Two explanations compete for why graph neural networks misclassify heterophilic nodes: that they form a rare subgroup sacrificed for majority accuracy, or that neighbourhood aggregation itself corrupts their representations. Evaluating six GNN architectures on five datasets of varying homophily, the authors find homophilic nodes are easier to classify even when they are the rare group, which undercuts the subgroup-generalisation framing. They also qualify the aggregation story: much of the information needed to classify heterophilic nodes correctly can be recovered by retraining just the classification head, or even only the final linear layer.
EMASAM: a Computationally Efficient Sharpness-Aware Minimization via EMA-Guided Perturbations
Sharpness-Aware Minimization (SAM) improves generalization by first perturbing weights toward a worst-case direction, which requires an extra gradient computation and roughly doubles training cost. EMASAM drops that gradient entirely, taking the perturbation direction from the discrepancy between the current model and its exponential moving average shadow copy, which pushes away from the stable averaged position as a softer stand-in for the worst-case step. Because the direction no longer depends on noisy mini-batch gradients, it also avoids the gradient-induced instability of SAM, and experiments report SAM-level generalization without the second backpropagation pass.
MiNO: Cotangent-bundle propagator learning for PDEs
Machine learning for partial differential equations usually fits either the solution field (physics-informed neural networks) or the solution map (neural operators), but a transported discontinuity is nonsmooth even when the rule that moves it is a smooth polynomial phase. The microlocal neural operator (MiNO) targets that propagator directly, learning a phase from the eikonal equation and an amplitude from the transport equation and reconstructing solutions via an oscillatory integral, so sharp fronts and caustics come from propagation geometry rather than pointwise fitting. On a matched-budget discontinuous-advection benchmark MiNO reaches the closed-form accuracy limit of its reconstruction window within 10,000 steps while a physics-informed network with neural-tangent-kernel loss balancing barely improves; on smooth advection its mean error is 3.84×10⁻³ against 3.12×10⁻² for a supervised Fourier neural operator, and one trained generator transfers to five unseen initial conditions without retraining.
Measuring Structured Predictability in Neural Training Dynamics: A Cross-Regime Study
Training trajectories are rarely characterized as objects in their own right, so this work measures short-horizon predictability — how much recent parameter updates tell you about near-future parameter motion — using three probe families (displacement-direction, subspace-residual, and predictor-based) with null-calibrated, group-level readouts. Applied to multi-pass vision training on CIFAR and to public Pythia pretraining checkpoints, the probes agree with each other and with independent trajectory diagnostics, suggesting they capture real structure rather than probe artifacts. The consistent pattern is that vector-like parameters such as normalization scales and biases move far more predictably than matrix-like feature-transforming weights, whose predictability instead concentrates in localized, shifting pockets; a Pythia-70M case study traces role-, depth-, and scale-dependent events including predictable query-key-value pockets migrating across layers.
Learning Auditable Classifier Models: Source-Disjoint Tree Ensembles
Clinical and regulated deployments need models that are both accurate on tabular data and auditable one prediction at a time, but boosted ensembles entangle structure discovery with coefficient fitting and interpretable alternatives restrict interactions or emit overlapping rules. Residual Pattern Tree Ensemble (RPTE) runs three stages: build a supervised symbolic feature vocabulary, grow shallow trees under a source-disjointness constraint that assigns each raw variable to at most one tree, then fit a single L1-regularized logistic regression over leaf-region indicators, so every prediction decomposes into a sum of named, non-overlapping rule contributions. Across twelve clinical binary classification benchmarks with repeated stratified 5-fold cross-validation it stays competitive with tuned opaque ensembles, and cuts model inspection units by 9x to 87x relative to XGBoost while beating EBM on audit complexity for all twelve datasets.
FirstDiff: One-Step Diffusion-Based Anomaly Detection for Multivariate Time Series via Initial Noise Prediction
Diffusion models detect anomalies in multivariate time series by learning what normal data looks like, but existing methods score anomalies only after running the full reverse denoising trajectory, which is expensive and discards the intermediate signals. The central observation behind FirstDiff is that the noise predicted at the very first reverse-diffusion evaluation already carries enough information to detect anomalies, so the method fits a statistical distribution over that predicted noise on normal validation data and scores new windows from a single network evaluation. A Diffusion Transformer serves as the denoising backbone to capture temporal and cross-sensor dependencies, and across five public benchmarks the approach reports state-of-the-art detection at a fraction of the inference cost.
Geometry of Forgetting: Representation Flux in Continual Learning
Work on catastrophic forgetting mostly attacks it through parameter regularization or replay, leaving what happens in representation space during sequential learning less examined. The measure introduced, representation flux, tracks sample-level displacement of latent representations across training and turns out to be strongly associated with forgetting across benchmarks, with elevated flux appearing before performance degradation; displacement also tracks confidence degradation. Building on that, FlowLess-R adds an architecture-agnostic representation-matching term that pins replay representations near stored references while still permitting learning, and it improves final average accuracy and reduces forgetting on SplitMNIST, SplitFashionMNIST, SplitCIFAR10 and SplitTinyImageNet when combined with ER, DER++ and ER-ACE.
Breaking the Compression Barrier: Cross-Architecture Compression Boundary Learning via Reverse Regrowth
Pruning shrinks networks for edge deployment but tends to collapse abruptly past a sparsity threshold, making the true compression limit hard to locate. BRIDGE inverts the usual procedure: it first pushes the model into an extremely sparse, collapsed state to expose the failure region, then selectively regrows structure using coarse-grained layer selection followed by fine-grained choice of which parameters to restore. On both convolutional networks and transformers, the approach recovers models from near-collapse, yielding up to 4.77% higher accuracy under structured pruning and up to 1.49% under unstructured pruning.
Unifying Graph Neural Networks Through a Common Layer Equation
Graph neural network families are usually written with family-specific notation that hides which computations they share and where they genuinely differ. The proposed common layer equation expresses architectures through seven slots — update domain, channel set, propagation bank, per-channel message maps, channel-fusion operator, ego/residual map, and update map — with the central factorization separating where information moves (the propagation bank) from what moves (the message maps), and function-valued fillings extending it across local message passing, attention, spectral filtering, global communication, relational channels, higher-order domains, and geometric messages. Worked reductions of canonical layers make the unification checkable, and component-level results follow: under endpoint-local messages and node-local updates, operator support bounds one-layer dependencies, and one-layer global mixing requires a full effective operator row. The framework places more than 200 architectures in one design space, supports component-wise comparison and generation of new consistent architectures, and links propagation choices to oversmoothing, oversquashing, heterophily, and expressivity.
Efficient Coreset Selection via K-Nearest Neighbor Graphs
Coreset selection shrinks a training set to a small representative subset, but gradient-approximation methods such as CRAIG and its clustering variants need dense pairwise distance or item-cluster bound matrices, which becomes prohibitive at scale. KNNG-CS builds a K-nearest-neighbor graph, scores each item's importance from its local neighborhood, and greedily picks representative nodes, requiring storage only linear in the number of edges. On four real-world datasets it matches the accuracy of gradient-approximation baselines while cutting selection time by 2.3x to 41.2x and peak memory to between 0.3% and 7.5% of theirs.
Localized TabICLv2: Scaling Tabular In-Context Learning through k-NN
Tabular in-context learning models such as TabICLv2 place the entire training set in context, so attention cost grows with dataset size and limits practical throughput. The localized variant retrieves only the k nearest training rows for each test point, measured by similarity in the model's own Stage 2 row-representation space, requiring no architectural change, with optional Stage 2 and Stage 3 fine-tuning to recover the accuracy lost by truncating context. On TabArena classification tasks the fine-tuned localized model keeps 98.64% of full-context accuracy while delivering a median 2.18× batch-inference speedup and roughly 249× in single-query serving.
A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not
Analyses of the Voynich manuscript almost always assume its glyphs are letters, its blank-delimited strings are words, and its blanks are word spaces; all three are tested here against the Zandbergen-Landini transliteration with matched prose, cipher, and pseudo-text controls plus quire-level resampling, and none of the three survives. Glyph regularity is too high for one-to-one substitution with any tested plaintext (conditional entropy 2.7 bits versus roughly 3.5 for Latin, Italian, and English) and instead resolves onto recurrent multi-symbol units; token identity predicts the next token by under 1% of token entropy, below every control at 2-10%, while glyphs at token edges share 0.2 bits of mutual information, above every prose control. Blanks split into two regimes, with transcriber-flagged uncertain separators behaving like word-internal junctures and measuring physically narrower on the page, and while a published imitation cipher and a self-citation generator both reproduce the low entropy and weak token order, neither reproduces the edge-glyph coupling or the hapax-rich vocabulary (70% singleton types against 41% and 59-60%).
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models
Time series foundation models are usually judged on fixed historical test windows, which measures average accuracy on a frozen past rather than behavior under seasonal drift, distribution shift, and unexpected events. LiveHouse-TS is a continuously running benchmark that evaluates models prequentially against real future data as it arrives, treating benchmarking as ongoing temporal validity instead of a one-off leaderboard. Streaming evaluations across 11 domains and 17 datasets found that static leaderboard rankings reshuffle dramatically under the live protocol.
40 more specialized papers
- ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning Harshil Lodhiya
- BRAID: Learning Equilibrium Maps in Interdependent Security Games via Weight-Tied Iterative Graph Neural Networks Elnaz Nowrouzi, Zhiqun Zuo, Xueru Zhang et al.
- Can Neural Networks Learn by Experimenting on Themselves? Self-Interventional Learning from Functional Consequences to Predictive Self-Knowledge Micha{\l} Tomaszewski
- Degeneracy Counting Quantum Algorithm using Decoherence Malay Marut Das, Mark A. Novotny, Yaroslav Koshka
- GATTA: Graph Active Learning with Test-Time Augmentation Zsombor B\'anfi, Andr\'as G\'ezsi, Andr\'as Formanek
- Global Federated Learning Strategies for Building Efficient Personalized Models Seongyoon Kim
- An Adaptive Gradient Clipping and Noise Injection Mechanism for Differentially Private Federated Learning Wenjing Wei, Alla Jammine, Farid Nait-Abdesselam
- Identifying parameter couplings and uncertainties of mixed-noise stochastic systems via full-covariance Gaussian mixture network Xiaolong Wang, Xiangwen Hao, Jing Feng et al.
- Decentralized Federated Learning for Heterogeneous Multi-Task Semantic Communication Lin Yin, Tiejun Lv, Weicai Li et al.
- FedADB: Class Anchor-Driven Dual-Branch Federated Learning for Mitigating Forgetting Zhenyan Liu, Hua Zhang, Haoran Gao et al.
- Shape Operator PCA: Curvature-Aware Projections for Geometric Machine Learning Alexandre L. M. Levada
- Beyond Field Accuracy: Two-Axis Diagnosis of Inverse-PINN Parameter Error Yifan Zhang, Qian Tao
- Look Before You Lift: Visual and Quantitative Diagnostics for Topological Deep Learning Mathilde Papillon, Guillermo Bern\'ardez, \'Alvaro Ball\'on Barreiro et al.
- FAST-DeepONet: Factor-Augmented Branch Representations for High-Dimensional PDE Inputs in the Small-Sample Regime Jiyong Kwon, Bongseok Kim, Guang Lin
- QSMP: finding representative time series subsequences through Quick Shift+Matrix Profile Carlos H. Mendoza-Cardenas, Rogers F. Silva, Austin J. Brockmeier
- Quantum Models with Multi-Stage Training for Compositional Concept Generalization Mina Abbaszadeh, Matilda Karabina Moore, Mehrnoosh Sadrzadeh et al.
- When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation Wenhao Yuan, Chenchen Lin, Wenhao Hu et al.
- Adaptive Heterogeneous Compression for Resource-Efficient Federated Knowledge Distillation Chenwang Liu, Yijun Liu, Chang Liu et al.
- QuantumPhaseNet: A Gauge-Covariant Geometric and Quantum-Spectral Theory of Semantic Concept Hierarchies with Prototype Validation of a Classical Quantum-Inspired Model Kiyotaka Kasubuchi, Kazuo Fukiya
- RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection Junxin Lu, Jing Zhao, Shiliang Sun
- NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption Ziluowen Luo, Jun Yin, Ruochen Liu et al.
- Task-Anchored Representation Shaping for Pre-Trained Model-Based Continual Learning Zhiming Xu, Huiyu Yi, Zhen-Hao Xie et al.
- SoftModel: A Neural Model That Grows Its Own Topology -- Governed Structural Growth for Continual In-Service Learning Zhoumin Xie
- Variational Outlier-Robust Gaussian Process Regression with Generative Modeling Arslan Majal, Aamir Hussain Chughtai
- Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning Daniel Nowak Assis, Jean Paul Barddal, Fabr\'icio Enembreck
- Hide&Seek: Learning to Explain in an End-to-End Differentiable Network Tal Ellinson, Hadi Mohasel Afshar, Sally Cripps
- Beyond $L_2$: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures Jules Soria, Alban Grastien, Romain Xu-Darme et al.
- An Investigation of the NeurIPS and ICML 2025 Position Tracks Fan Yang, Wenkai Li, Jun Liu
- What If AI Carried Her Imagination? Black Girls as Creators in an AI Storytelling Weekend Program Chun Li, Lauren Brown, Hubert Asare et al.
- AI, Brain Death Detection, and Islamic Law Muhammad Aurangzeb Ahmad
- The politics of postmortem privacy Mauricio Figueroa
- ComNetX: Local Hierarchical Adaptation for Dynamic Community Detection Aleksandr Konovalov, Anna Uporova, Alexander Drobyshev et al.
- Education-centered critical policy analysis of AI: Ghana's AI strategy as a case Matthew Nyaaba, Vida Awinime Bugri, Eric Kojo Majialuwe et al.
- Average Distance Approximation for Static Large Graphs Kartikey Ahlawat
- EMAN: Optimization-Driven Capacity Growth through Path Emergence in Multi-Task Learning Chenlei Fang, Jingchen Li, Hongzong LI et al.
- Without journalists, there is no journalism: the social dimension of generative artificial intelligence in the media Sim\'on Pe\~na-Fern\'andez, Koldobika Meso-Ayerdi, Ainara Larrondo-Ureta et al.
- From Abductive Explanations to Global Logical Rules for Node Classification in SGCs Bryan Lima Cavalcante, Thiago Alves Rocha
- Rethinking Irregular Time Series Forecasting from the Perspective of Basis Functions Rongwen Li, Changjian Chen
- Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting Rongwen Li, Haixin Xie, Xiao Wang et al.
- SPACE: Sample-cloud Predictive Adaptive Conformal Ellipsoids for Multivariate Time-Series Forecasting Baishi Li, Kelvin J. L. Koa, Ke-Wei Huang
Agents 47
WANDR: A Benchmark for Wide and Deep Research
WANDR is a benchmark of 500 realistic data-collection tasks that require a research agent to find many entities matching stated criteria (breadth), investigate each through multiple coordinated web searches (depth), and return independently verifiable records with sources and excerpts. Tasks are encoded as qualification key hierarchies — n companies by m employees by k sources means n×m×k records — covering workflows such as market mapping, due diligence, literature review, and talent sourcing, with targets from dozens to thousands of records. Rather than static gold answers, task-specific judges refetch each cited page and check the record against its evidence, aggregating into soft and hard precision, recall, and F1 so that changing facts remain gradeable. Across six production research systems the best reaches only 0.363 soft F1 and 0.133 hard F1 at high effort, with performance falling as target volume and hierarchy depth grow, and incomplete discovery, missing enrichment, and thin evidence as the main failure modes.
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks
Multi-agent large language model systems buy accuracy with extra computation, so the practical question is deciding per problem whether the extra collaboration is worth its token cost. Holding the solver fixed, every problem is run under four protocols — direct solving, iterative self-correction, planner-executor-reviewer, and multi-agent deliberation — on 4,181 competition math problems plus paired checks across math, biology and broader science with two solver families, then compared against fixed policies, trained routers and frozen LLM routers. A post-answer gpt-oss-120b probe ranks baseline failures well (0.8847 AUROC) and still predicts whether any collaboration helps (0.7683 AUPRC), but collapses when asked which protocol pays off, at 0.1674 and 0.1041 AUPRC for planner-executor-reviewer and deliberation respectively. A pre-answer self-confidence gate hits 78.0% solve at 45K tokens versus 73.8% at 71.3K for the frozen router, while a retrospective fixed-order oracle reaches 92.4% and retains 18.5-28.9 point gaps over learned routers.
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads
Peak GPU memory for LLM inference is often modeled analytically as weight storage plus key-value cache plus activations, parameterized by step count, tool invocations, and context growth, and the question is whether such models pay off for agentic workloads. The authors measure 1,920 trajectories of AgentK, a LangGraph CUDA-kernel-synthesis agent, running four 4-bit-quantized backbones on a single H100. Closed-form models matched or beat the best learned baseline on three of four backbones, but the more consequential finding is that peak memory barely varies at all in this regime (coefficient of variation 0.3–9.4%), so learned prompt-feature regression gives no statistically significant improvement over predicting a constant and complex VRAM predictors are unjustified. Compile success instead split sharply by model capacity, from 5.7% for Phi-4-mini to 62.0% for Qwen2.5-Coder-14B, indicating code synthesis is limited by the model rather than by available memory.
No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage
Audits of large language model agents usually grade a single run per task using transcripts or a judge model, which the authors argue systematically understates the damage agents cause. AgentRelBench measures ground-truth, severity-priced damage from database state diffs across repeated runs with no model in the measurement path, demonstrated on EnterpriseOps-Gym over 2,128 runs and nine models. Damage on irreversible actions appeared in every model family measured but was stochastic within them: no task damaged on every run, so a single clean run misses a damage-producing model-task pair 80% of the time on the development pool, and for the most capable frontier model a single audit misses its one damaging task 84% of the time. One family committed a gated irreversible change while its transcript declared a refusal, which transcript- and judge-based grading scored as safe.
ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems
Clinical multi-agent systems built on language models raise safety, fairness, accountability, and transparency concerns that existing healthcare AI ethics frameworks state as principles rather than running code. ETHOS (Ethics and Trust through Hierarchical Oversight System) is a governance meta-agent that attaches to an existing multi-agent system without architectural changes, applying layered runtime oversight — deterministic checks, then contextual reviews, then a final ethics critic — to intermediate reasoning steps and final outputs so it can flag risks, demand revisions, or suppress unsupported responses. Demonstrated on a hepatology decision-support system, it improves decision reliability largely by detecting incomplete, inconsistent, or out-of-scope evidence and abstaining more often when no safe recommendation is supportable.
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
Scientific search over molecules, protein sequences, or programs means optimizing expensive objectives across open-ended spaces where language model likelihoods and self-assessments are poorly calibrated proxies, especially off-distribution. The Large Discovery Model couples a generative proposer with a Bayesian non-parametric reward surrogate: the surrogate predicts performance and its uncertainty, that uncertainty-aware value steers generation, refinement, and selection, and both the surrogate and a discovery memory update as each experimental result arrives. Tested on neural network training, antibody design, and molecular optimization against language-model-only reflection and classical statistical search, it achieves a 2.4x larger reduction in validation bits-per-byte, an 18.2% relative drop in binding energy, and over 60% relative gains on multi-objective molecular tasks.
CoupVisor: Strategy Optimization by Round and Challenge Decision Support
A decision-support system for the hidden-information card game Coup tackles two questions a player faces each turn: which action to take, and whether to challenge an opponent's claimed role. Everything — manual play, replay of recorded games, simulation, belief tracking, advisor recommendations, and learned policies — runs off a single shared description of game events, and truthfulness estimates combine each role's prior likelihood with how many cards the claimant still holds, which fixes a case where the opening claim of a game was wrongly flagged as suspicious. Comparing a rule-following advisor against several learned and heuristic players over many simulated games and opponent styles, the reward definition decides which learning approach wins, and a reward tied to eventually winning the game produces a policy that beats every baseline.
Deploying Frontier Agentic Technology in MOOSEnger, a Multiphysics-Capable AI Assistant
Setting up simulations in MOOSE (Multiphysics Object-Oriented Simulation Environment), an open-source finite-element framework, requires specialist expertise that blocks many domain scientists. MOOSEnger is a tool-enabled agent for MOOSE, extended here with a harness aimed at locally hosted models: it retrieves context from the MOOSE repository, validates and diagnoses generated input by actually running the simulation executable, and stores extracted lessons in persistent memory. Evaluated on eight problem categories of 25 cases each — diffusion, Navier-Stokes, phase field, plasticity, porous media flow, solid mechanics, transient heat transfer, and reactor mesh generation — the agentic versions reach 90% success with GPT-5.2 and 76.5% with Gemma4, while the same models without the harness score just 5% and 0%.
Crystal-structure design by agentic AI in a language of motifs
Data-driven materials discovery tends to interpolate within known structure types rather than reach genuinely new ones. MatEvolve is an agentic framework that proposes each candidate crystal with a stated rationale and then tests it, reasoning in an interpretable language of motifs: every crystal is written as a motif profile describing its recurring geometric patterns, and the agent designs by editing that profile and rebuilding a crystal from it, with promising candidates checked by first-principles calculation. Applied to rare-earth-lean permanent magnets and built on Claude Fable 5 without fine-tuning, it reaches new structural prototypes more than three times as often as generative models under an equal validation budget at a comparable on-target rate, and the human-readable profiles of discovered crystals expose structure-property relationships.
Mint-Agent: Introducing Finance-Native Agentic Foundation Models
Financial agents need both precise grounded operations and the stamina to sustain long research trajectories whose evidence remains auditable, and Mint-Agent targets those two axes with a data engine for atomic and long-horizon financial tasks, an interaction harness that maintains evidence trails, and a training recipe combining supervised fine-tuning, critical-step preference optimization, and reinforcement learning with verifiable rewards. Separate reasoning and execution experts are unified via model merging and multi-teacher on-policy distillation into Mint-Cu (9B) and Mint-Ag (27B). The larger model reports 98.33% on RFC-Bench, 3.66 points above GPT-5.6-Sol and 3.00 above Claude-Opus-4.8, while Mint-Cu reaches 69.86% on FinSearchComp T2, ahead of Agents-A1-35B by 22.83 points.
When Tool-Backed Skill Retrieval Fails: Source-Style Collapse in Executable Capability Retrieval
Before an agent can plan with or invoke an external capability, a retriever has to surface it, and this retrieval gate can fail silently even when the tool corpus never changes: on ToolRet, a retriever fine-tuned on one source-specific slice collapses on a different slice of the same benchmark despite high lexical overlap with the gold tools, a failure the authors name source-style collapse. Query-side TF-IDF fingerprints predict which source styles will trigger the collapse better than semantic or length-based proxies, and ToolScout turns that cheap signal into a source-aware routing guard. On a mixed 4,996-query stream routing raises coverage from 22.3% to 86.1%, and the same failure and repair persist when tools are rewritten as executable skill cards, ruling out raw API schema formatting as the sole cause.
The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks
Repository-scale coding is modeled as reconstructing a coupled-fact graph, where each edit needs a fact that comes either from recent context or from the model's parametric memory — facts supplied by neither are termed coherence debt. Across seven models and five harnesses, each channel is systematically supplied or withheld: no model completes a task on an unseen API when both are empty, a rename that invalidates memorized library knowledge makes all seven models fail identically, passing and missing the same tests, and a supplied fact helps as much far from the edit as adjacent to it, so availability rather than distance determines success. Harnesses that all pass every test differ more than tenfold in tokens because they re-fetch the same content at different rates, and a withheld fact yields fabricated files and guessed values instead of a refusal, so read-based instrumentation sees a hole already filled. Agents also follow a stale convention file over the actual code, making a wrong convention file worse than none, and on SWE-bench — where repositories are likely memorized — read behavior no longer predicts success.
ClawGym II: Exploring Black-Box RL on Agent Harness
Agent harnesses coordinate a model's interaction with its environment on long-horizon tasks, but training the underlying model with reinforcement learning (RL) through such opaque harnesses has largely been avoided because of scaling and attribution problems. The proposed black-box RL framework isolates each task environment and harness in a temporary sandbox for large-scale concurrent rollouts, inserts a serving proxy at the model boundary to capture every model call, reconstructs multi-turn trajectories as prefix trees, and adapts both critic-based PPO and critic-free GRPO to optimize over that tree structure; mix-harness training lets one model be optimized jointly by heterogeneous harnesses. With Qwen3-30A3B, the method raises Pass@1 on ClawGym-Bench by 9.98 points through OpenClaw and 14.81 points through Claude Code, staying stable across 200-400 optimization steps, with further gains on JobBench and OfficeQA.
AutoSR: Automatic Symbolic Regression by Searching Research States
Symbolic regression systems search for equations, but finite noisy data admit many numerically competitive expressions that extrapolate very differently, so fit and syntactic complexity alone are weak evidence of scientific credibility — and conventional searches discard the motivations and probes that would justify a choice. AutoSR searches persistent investigations instead of isolated equations: each candidate is stored in a Research State bundling the equation with its reasoning, computational evidence, and independent review, developed by proposer-reviewer agents under progressive-widening Monte Carlo tree search and finally synthesized into a report defending the selected relation. Across nine challenges from two benchmark suites it recovers an algebraically equivalent relation in every case, including three cp3-bench problems no published system had solved plus six structurally diverse LSR-Transform problems.
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents
Turning clinical trial protocols into analysis-ready datasets under CDISC standards gates regulatory submissions, and it defeats direct code generation: across 11 single-shot attempts with five frontier models, none produced a valid subject-level analysis dataset. GxP-Agent encodes the regulatory process ordering as a directed acyclic graph, splitting dataset generation into 15 domain-specific nodes run by worker agents with pharmaverse skill context, validation gates, and conditional retry. On CDISC-Bench, built from the FDA pilot submission CDISCPilot01, it reaches 100% structural match on all 49 ADSL variables and 254 records across three independent runs, against 59.2% for the best retrieval-augmented baseline and 0% for every single-agent and flat multi-agent setup; the same topology lifts GPT-4.1 from 0% to 59.2%, and the adverse-events dataset is matched exactly on the first attempt.
CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents
Large language models prompted as agents can produce plausible city-scale daily routines, but few-shot prompting tends to reproduce the model's own behavioral priors rather than the statistics of the population being simulated. CityReal instead treats each agent as an intention-driven decision maker that pursues coherent mobility and activity plans, learns habits and preferences from experience and constraints, and whose behavior modules are tuned by learned textual adapters that align decisions with observed population data. Alignment with real human behavior improves at both the individual and population level, and the framework scales to tens of thousands of agents to analyze crowd density, place popularity, mobility flows, and well-being under different urban scenarios.
QuantumNovelty: A Skill-Orchestrating Language Agent for Referee-Style Review and Patentability Screening of Quantum Papers and Patents
Given that language-model agents now generate quantum-science artifacts, QuantumNovelty asks whether the same agentic setup can also scrutinize them, combining generation of papers, Pareto-front ansatz candidates, and patent drafts with simulated referee and patent-examiner panels. Its design contribution is an audit-and-falsify layer of deterministic gates — strict Pareto domination checks, numerical recomputation from stored artifacts, Wilson small-sample intervals, and a cross-vendor consensus guard — that restricts which claims survive rather than producing them, with every model call logged by backend, token count, and cost. On a planted adversarial corpus the deterministic gates caught every seeded overclaim with no false positives, and a first deployment over six manuscripts and one granted patent cost roughly twenty-four US dollars, with the panels running more conservative than the public acceptance record on a one-sided sample. The authors make no accuracy claim against human experts and document which mechanisms remain untested on real inputs.
SeqFeed: Improving Agentic RTL Code Generation with Sequential Behavior Feedback
Agents writing register-transfer level hardware code must reason about how signals evolve across clock cycles, but the source itself does not expose cycle-level behavior for a given execution and full simulation waveforms are too large and noisy for a language model to read. Studying how human engineers debug sequential behavior yields three requirements for useful feedback — it should be event-addressable, dependency-traceable, and iteratively queryable — which SeqFeed meets with two mechanisms: SeQuery, an SQL-like waveform query language that anchors queries to semantic events and samples signal values at relative times, and SeGraph, a dependency graph tracking signal propagation across cycles. Pass rates improve across multiple language models, with each mechanism effective on its own and the two providing complementary gains when combined.
SkillEffect: Checked Lowering for Memory-Bounded Agent Tools
Agent Skills describe how a tool should be used, but when a model turns that guidance into code, a semantically correct program can still stream an entire input into memory and blow the budget for a single tool call. SkillEffect is a runtime that only grants execution authority after an independent checker re-derives the proposed "lowering" — a bounded implementation of the requested computation — from the submitted program and the immutable input, with each supported operator family supplied by an audited plugin providing a recognizer, bounded intermediate representation, arena bound, and output postcondition. Across six operator families and five execution patterns, bounded access substantially cut peak memory and raised completion rates under externally fixed memory caps, with the checker accepting all legal configurations tested and rejecting all adversarial ones; generality is explicitly architectural, since each new computation still needs a hand-audited plugin.
KernelArc: A Multi-Agent Framework for GPU Kernel Optimization
KernelArc tackles GPU kernel optimization by running strategy-specialized agents in parallel, coordinating them through a conclusions-only shared memory, a deterministic benchmarking guard, and read-only access to each other's state with drafting triggered when an agent plateaus. Evaluated on NVIDIA H100 and B200 hardware against category-representative SOL-ExecBench workloads, the system produced custom BF16 GEMM kernels, cuBLASLt Expert-API configuration tables, fused mixture-of-experts backward passes, shape-gated decoder-layer fusion, native NVFP4 grouped-query attention, and paged prefill attention. Those submissions ranked first on representative L1, L2, Quantization, and FlashInfer tasks at the July 30, 2026 leaderboard snapshot, though the authors note the value of any individual coordination mechanism varies by kernel and optimization stage.
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection
Choosing the best solver for a constraint satisfaction instance depends on features capturing problem structure, and hand-designing those extractors requires deep domain expertise that becomes a bottleneck whenever a new problem class appears. The approach places an LLM in an agentic check-fix-verify loop that synthesizes executable Python feature extractors from a high-level MiniZinc model and an instance, constructing a typed graph and computing properties such as graph density, variable clustering, and constraint tightness. Across vehicle routing, car sequencing, and fixed-length error-correcting codes with a five-solver portfolio, the synthesized and still-inspectable extractors beat expert-curated mzn2feat features by up to 8.3 percentage points of test-set accuracy and also beat transformer-based trans2feat variants.
Token Optimization and Context Window Management in Multi-Agent AI Workflows
Grounded in a production dashboard that extracts structured work items from meetings, email, and chat, six practitioner patterns are described for reducing token cost and latency in multi-agent workflows: context stratification, fetch-once/process-locally architecture, schema-contracted prompts, token-aware fallback chains, semantic caching, and inter-agent communication compression. In production these cut measured cold-load latency to 61-116 seconds from an operational baseline of roughly 3.5-10.5 minutes, with an estimated 60-70% token reduction. A separate controlled study of 2,420 trials across 11 model configurations found that replacing some high-relevance items in a fixed ten-item prompt with same-domain low-relevance items improved relevance-score accuracy by +0.077 (Cohen's d = 0.49), an effect the authors call relevance-contrast context and report as a within-corpus descriptive comparison rather than a population inference. A Fusion-of-N follow-up found learned synthesis no better than a mechanical set union of item IDs.
Graphectory Viewer: A Tool for Process-Centric Analysis of Agentic Software Trajectories
Understanding why a software agent succeeded or failed usually means reading raw execution traces, which obscures higher-level behavioral structure. Graphectory Viewer is a web-based tool that converts trajectories from multiple agent frameworks into phase-aware graphs, offering node-level inspection of thoughts, actions, and observations, search and filtering across large collections, and Sankey-style summaries of problem-solving phase transitions. It ships as an open-source artifact together with documentation, precomputed graphs, and a large-scale trajectory corpus for comparing successful and failed runs at scale.
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification
Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects the current task, and the authors turn that decision into an audit protocol for structured intermediate outputs: test for dataset shortcuts, unbundle prompt changes, check whether intermediate labels are merely answer-associated, probe decomposed semantic evidence, and record provider-side execution failures. A 480-example synthetic development set initially suggested large gains from a state-structured prompt bundle, but TF-IDF diagnostics exposed lexical separability, prompting a frozen 160-example counterfactual set with 40 matched four-way families and rule-derived reference policies. On that set, exposing the four state definitions helps, but an isolated explicit state-output field yields no significant gain for Llama-3.3-70B and only a marginal one for GPT-OSS-120B, and complete four-way family success is rare, showing example-level accuracy overstates counterfactual consistency. Only policy classification is evaluated, not downstream responses, tool actions, or actual memory-store mutation.
ASI-Bench: At the Dawn of Artificial Superintelligence
Current benchmarks largely test whether AI can produce correct answers from learned knowledge or complete tasks under extensive human guidance, leaving open how far systems get when that guidance is withdrawn. ASI-Bench offers 60 project-level research tasks across 11 scientific domains, built by over 40 experts at a cost of 31,000+ human hours, and progressively removes methodological guidance within the same research project so a system must eventually pick its own method, run the work, and produce verifiable results, with all tasks passing expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 agent-model configurations the average score drops from 50.91 with full methodological guidance to 29.10 when only the method is specified and 26.62 when agents must determine the method themselves, which the authors read as continued heavy dependence on human direction. The benchmark is open for outside task contributions.
When Agents Act on Web3: An Attack-Surface Survey of MCP, Skills, and Tool Calling
Agents built on the Model Context Protocol (MCP) increasingly write rather than read — the share of deployed tool use that modifies external state has risen from 27% to 65% — and when that authority reaches a public blockchain, four properties of the execution layer (irreversibility, signing authority, continuous autonomy, and sequence-level composition) turn ordinarily recoverable agent failures into standing, irreversible loss. The survey organizes scattered MCP-security work into an attack-surface taxonomy and adds a Web3 risk-mapping matrix tying each attack class to its amplified impact, the responsible amplifier, a representative mitigation, and the residual gap. Reviewing defenses including emerging blockchain-based mechanisms, it finds them improving but insufficient: measured protections stop fewer than 30% of attacks and model-level safety refuses fewer than 3%, with the matrix's empty cells forming the proposed research agenda.
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation
Multi-agent reasoning systems usually route work through a central protocol, creating bottlenecks and static role assignments that hold up poorly on complex multimodal queries. DeAR replaces central control with autonomous peer-to-peer collaboration built on three mechanisms: decentralized capability grounding so agents specialize per query, thought map navigation to target which peers to consult, and topology updates for adaptive error correction. Across nine multimodal reasoning and text question-answering benchmarks it consistently outperforms recent baseline methods, which the authors take as evidence that decentralized, adaptive collaboration helps on knowledge-intensive reasoning.
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs
Group-relative reinforcement learning methods for multi-turn agents give every successful rollout the same outcome reward, so a trajectory that wanders for twenty turns is indistinguishable from an efficiently planned one — a collapse of the advantage signal that caps performance. PlanPO adds coarse-to-fine advantage terms that compare trajectory length and per-turn response length among the successful rollouts sampled for the same task, teaching deliberate interaction planning without degenerating into simply minimizing length. Averaged over ALFWorld, WebShop, and SciWorld, it reports a 27.2% gain over GRPO at negligible extra training cost.
Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents
Browser agents that look strong on short scripted demos break down on live sites, where success requires dozens of chained decisions, error recovery, and messy user interfaces. Wuying-Browser-Agent aligns the whole stack rather than just scaling: a structured harness supplying stable execution primitives and decision-oriented context, curriculum fine-tuning on recovery trajectories and complex-UI episodes (RUIC-SFT), and an online reinforcement learning stage with potential-based reward shaping and divergence-aware step weighting (DAO-GRPO). The accompanying BrowserBench holds 350 bilingual real-web tasks averaging 37.9 steps, and the 27B model sets a new open-source state of the art with 80.6% on WebVoyager, 66.7% on Online-Mind2Web, and 65.1% on BrowserBench, with the pipeline also transferring to general agentic suites at an average of 73.8.
Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations
LLM agents running mission-critical infrastructure are typically handed the same comprehensive harness — all context, all tools, all permitted actions — regardless of the task, which wastes resources. Treating harness selection as resource matching, tasks are classified by the mathematical representation of the underlying system and harness configurations ranked by how much and what kind of information they provide, with task-to-harness mappings drawn both from mining research literature and from controlled agent runs. The resulting map-guided escalation algorithm starts from a task-specific harness and expands to full provisioning only after a failed self-check: in liquid cooling it raised accuracy from 0.652 under full provisioning to 0.715, matching Reflexion with 48% fewer tokens, while in power grids full provisioning stayed the most accurate and the map only offered cheaper alternatives. The authors conclude that harness provisioning sits on a domain-dependent accuracy-cost Pareto frontier rather than having a universal optimum.
When AI Designs AI: Innovation or Imitation?
To test whether LLM agents that design machine learning methods are inventing or merely recombining, the authors derive task-specific algorithmic design spaces from human-designed methods and map both human and agent solutions into those spaces, quantifying differences at the module level. Widely used agents were evaluated on open-ended AI tasks spanning multiple modalities, matching or beating human state of the art in 10 of 72 configurations, though the wins did not generalize reliably across tasks or agents. 96.8% of agent-designed methods fell inside the human-derived design space, and nearly half exactly reproduced an existing human design, indicating current agents largely reuse and recombine known algorithmic choices.
SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
Automated 'AI scientist' systems lean on proprietary frontier models precisely at the stage where research problems are formulated, which makes the evidential basis opaque, exposes proposals to model-specific hallucinations and biases, and can send confidential research material to external APIs. The Structural Gap Hypothesis Agent (SGHA) works corpus-first instead: it structures a literature collection into evidence-linked paper objects and a typed evidence graph, detects unresolved structural patterns across papers, screens candidate gaps before formulating them, and emits traceable research-problem families with explicit assumptions, objectives, success criteria, and remaining ambiguities. Every LLM component runs on a locally served open-weight 9B model, and a comparison with the AI Scientist-v2 idea-formulation module across five machine learning domains suggests explicit corpus structure and evidence-constrained reasoning can support inspectable problem formulation without frontier-model access during generation or verification.
ArborMem: Navigating Interaction States with Memory Forests
Long-running assistant conversations interleave several tasks, people, and plans that get interrupted and later resumed, but most memory systems retrieve relevant past text without first deciding which prior interaction state the current turn is continuing. ArborMem represents the conversation as a navigable forest of interaction states where each branch holds a locally coherent trajectory: for each new input it localizes the relevant state, restores that branch's context, and augments it with reusable evidence retrieved from other branches, avoiding conflation of semantically related but structurally distinct threads. It also contributes BranchMemEval, a controlled diagnostic benchmark for interleaved and resumable trajectories, and beats the strongest baselines by 3.36 to 10.31 percentage points on LongMemEval, LoCoMo, and BEAM 100K, plus 5.0 points on the new benchmark, with the advantage widening under constrained read budgets and queries completing in under half a second.
Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback
Expert-written natural-language skills improve tool-using agents, yet agent-authored skills score 8-11 points below using no skill at all, suggesting that following procedural guidance and improving it from execution evidence are distinct capabilities. WER (Write, Execute, and Refine) trains a Skill Optimizer outside a frozen executor: the optimizer proposes a skill, the frozen agent executes it repeatedly, a programmatic verifier scores outcomes to give relative credit, and matched successful and failed trajectories from mixed-outcome records become the refinement states for the next training phase, so the optimizer learns from the consequences of its own earlier writing. On BFCL v4 multi-turn and tau2-bench it adds 7.80 and 3.85 Pass@1 points over the no-skill baseline and 9.35 and 10.29 points over the same backbone without optimizer training; the trained 4B optimizer reaches 76.63% on BFCL v4, beating every off-the-shelf general-purpose model evaluated in the same role.
TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation
Agent Skills bundle reusable natural-language procedures with executable resources, but judging a generated skill by its artifact or its final task score leaves open which actions the equipped agent will actually take and what side effects follow. TRUSS first checks functional claims against source and domain evidence while screening the whole artifact under nine safety properties, then loads survivors into a shadow agent inside a controllable execution environment where brokered tools subject requested actions to policy enforcement and record provenance-preserving traces; failures and violations are traced back to the responsible skill content to guide refinement. Across 168 SkillInject artifacts, 155 SkillSafetyBench cases, and all 187 SkillGenBench tasks it reports 100% precision and recall in vulnerability detection, cuts attack success from 38.71% to 19.35% with GPT 5.5 with no regressions, and raises task effectiveness from 17.11% without skills to 52.94% while lifting the benchmark security rate from 50.80% to 100%.
Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch
Self-evolving agents convert experience into reusable skills, workflows, or memories, and are usually judged on post-evolution accuracy alone, which says nothing about preserved correctness or security. This audit runs SkillOpt, Agent Workflow Memory (AWM), and ReasoningBank in a simulated e-banking environment using matched benign acquisition trajectories, sealed evaluation endpoints, execution-grounded checks, and independent state replay. SkillOpt lifts benign utility from 0.741 to 0.837 but its exposure to injected content rises from 0.820 to 0.943 and unauthorized financial state changes reach 0.685, while ReasoningBank gains utility without raising aggregate attack success. A separate hazard surfaced with AWM: its literal WebArena text-action envelope broke tool execution in a native function-calling executor, and removing only that envelope moved utility from 0.319 to 0.756 while attack success climbed from 0.195 to 0.575.
Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents
Oversight of long-horizon agents usually checks whether each individual action is locally valid, which misses trajectories that drift step by plausible step toward a broader role, an adjacent objective, or evidence the user never supplied. The authors define ontological trust as a task-conditioned property of trajectory prefixes and implement it as RGE, an online monitor decomposing trust into Role, Goal, and Evidence; language models only produce structured task and step representations, while trust updates, projections, and intervention decisions are deterministic, giving a replayable audit trail instead of a single judge verdict. On a corpus built from OSWorld, FinanceBench, and EICU-AC covering benign runs, prefix-paired drift, and pseudo-consistency failures, the two larger estimator models exceed 93% drift F1 on every benchmark while keeping benign coverage at or above 95.8%, though pseudo-consistency detection depends structurally on task completion being externally visible.
D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory
Persistent memory for LLM agents runs through ingestion, retrieval, filtering, and generation, so an end-to-end score reveals that something failed but not which stage caused it. D²ACCI is a dual-loop protocol whose outer diagnostic gate promotes, feature-flags, or rejects a memory intervention based on paired statistical comparisons, protected-slice non-regression monitoring, and trace-level localizability, quantified by a new graded observability metric called DCR. Instantiated in MemStack it scores 93.59% on LoCoMo, 90.93% on LongMemEval, and 57.20% on PersonaMem-V2, and paired ablations show supplement extraction, session-memory retrieval, and Forget Guard each give significant gains while hybrid BM25/RRF retrieval is retained only behind a monitored feature flag, a distinction aggregate-only evaluation cannot see. Enriched diagnostic traces reach 98–100% DCR@3 versus 0% for result-only logs.
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Most agent benchmarks use tasks chosen by researchers, which leaves open whether agent progress transfers to work people actually pay for. StartupBench derives its tasks instead from AI startup products with demonstrated market adoption, reconstructing their real user workflows as complete end-to-end deliverables scored against fine-grained rubrics. Under a unified agent harness, the strongest model evaluated completes only about 30% of the benchmark, though it makes substantial partial progress on many tasks; the dominant failure modes are complex instruction following and missing domain-specific expertise.
AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis
Agentic data analysis now runs for long stretches across parallel branches, and turn-by-turn chat interfaces give analysts neither a clear view of the agent's evolving reasoning nor a way to redirect it mid-run. AdaLens presents a storyline-based representation that unifies analytical plans, execution progress, intermediate findings, and which data columns are involved, with steering controls anchored to those same elements for directional guidance and execution control. Two case studies and a user study examine how analysts use it to monitor runs and cut off low-value directions or deepen promising ones while execution continues.
AutoResearch: Insight In, Hallucination Out
Autonomous research systems can now run long experimental workflows, but running them does not make the conclusions scientifically sound. AutoResearch splits the problem into Idea Generation, which fuses emerging research signals with accumulated domain knowledge and uses multi-model generation plus cross-review to produce testable plans, and Idea Execution, where coordinated agents decompose plans into experiments, iteratively debug them, and require independent evidence-based review before a conclusion is accepted. Across cross-modal retrieval, systems optimization, and benchmark-driven machine learning settings, a generated idea raised mean Recall on RSICD from 32.84 to 34.69 while logging only 5 audit-confirmed issue events versus 11–27 for competing autonomous research systems.
CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion
Long-term agent memory retrieval leans on semantic similarity, which recovers topically related content but misses the earlier plans, motivations, or experiences that explain a later event yet share little vocabulary with it. CABLE adds a sparse directed graph on top of an existing memory system: for each new memory it generates antecedent-oriented queries, retrieves prior memories, explicitly subtracts anything the host retriever would already find, verifies what remains, and links only those complementary associations, then expands retrieved seeds along those links at query time. Paired with A-MEM on LoCoMo and MA-LongMemEval and integrated into SimpleMem and Mem0g, it improves mean LLM-judge scores in every evaluated system-level configuration, with the largest gains on open-domain, multi-session, and preference-oriented questions.
EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection
Financial time series change so non-stationarily that no single unsupervised change-point detector works across assets and market regimes, leaving practitioners dependent on hand-tuned model selection. EvoTS-Agent first runs curated exploratory data analysis to characterize a dataset and seed candidate detectors, then evolves executable experiment trajectories through three operators — Revision to exploit the current best solution, Alternative Strategy to explore different modeling directions when progress stalls, and Recombination to merge complementary high-performing trajectories — with validation feedback steering the search. Across four benchmark datasets it outperforms existing LLM-based agents while maintaining a 100% execution success rate across all evaluated backbone LLMs.
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents
Agents that edit persistent artifacts such as spreadsheets, slide decks, PDFs, and notebooks can end up with the parsed text they search, the native file they edit, and the artifact they submit all referring to different versions of the same document. StagedWorkspace formalizes this as a workspace-state contract and enforces it by binding parsed records and review diffs to content hashes of the native files as they change, giving non-code artifacts the kind of search-diff-test guarantees repositories already give coding agents. In fixed-harness ablations, giving agents both parsed and native views beat either view alone by 8.3 to 12.1 points of Pass@1 on OfficeQA Pro and 4.7 to 9.2 points of mean rubric score on APEX-Agents, and a paired ablation over 57 file-editing tasks scored higher when diffs were visible.
Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating
Letting large language model agents converse on a user's behalf only works on a matching platform if people are willing both to send agent-mediated messages and to receive them, a two-sided condition that surveys of agentic recommenders rarely separate. Two large surveys of active users on a major dating platform (2,894 and 2,617 respondents, in two languages) were modeled with graded response models plus latent regression, showing that willingness to deploy an agent and willingness to engage someone else's are correlated at 0.92 yet statistically distinct. Deployment propensity runs roughly three times higher than engagement propensity, so under random pairing only 4 to 13 percent of directed pairs would have both a deployed agent and a receptive counterpart; routing agent contacts by receive-receptivity tripled per-contact engagement and held up out of sample at an area under the curve of 0.88.
On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
Agents that self-improve by accumulating a textual memory bank across a stream of tasks report steady gains, but those results typically come from single runs in one fixed task order. Re-running two such memory-based methods with multiple seeds and randomly shuffled task orders shows that evaluation noise in complex multi-step environments is already large and that the self-improvement loop amplifies it, while much of the reported gain depends on the default task ordering acting as a hidden curriculum. Inspecting the memories suggests task and environment underspecification is a contributor: adding detailed rubrics and environment feedback to memory construction recovers part of the lost performance but leaves significant gaps, and the authors argue for multi-run reporting, order stress tests, and interfaces that let humans oversee what agents write to memory.
1 more specialized paper
- Toward Personal Intelligence Through Cooperative Observation Yashar Talebirad, Osman Jime, Ali Parsaee et al.
Unclassified 40
Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis
No summary available — see the abstract on arXiv.
Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs
No summary available — see the abstract on arXiv.
Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews
No summary available — see the abstract on arXiv.
Learning Discrete Riemannian Metrics for Physical Fields with Cochain-Frame Equivarianc
No summary available — see the abstract on arXiv.
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL
No summary available — see the abstract on arXiv.
Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)
No summary available — see the abstract on arXiv.
Coarse-to-Fine Multi-Resolution Diffusion Models for Trajectory Generation in Urban Systems
No summary available — see the abstract on arXiv.
Longitudinal and Graph-Augmented Prediction of Adolescent Substance Use Onset in the ABCD Study
No summary available — see the abstract on arXiv.
Characterization of Thermal Systems from Noisy and Low-resolution Measurements Using Dynamic Mode Decomposition
No summary available — see the abstract on arXiv.
Evaluating the impact of adversarial traffic patterns on vanet communication using veins simulation
No summary available — see the abstract on arXiv.
6G Native AI and Channel Foundation Models
No summary available — see the abstract on arXiv.
Geometry Is Not Robustness: A Trajectory-Level Study of PGD Evaluation
No summary available — see the abstract on arXiv.
Demo: Real-time Generative Multicasting with On-Device Intent-aware Semantic Decomposition
No summary available — see the abstract on arXiv.
DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs
No summary available — see the abstract on arXiv.
Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion
No summary available — see the abstract on arXiv.
PIKFNO: An Interpretable Neural Operator Based on Physics Informed Kernel Function
No summary available — see the abstract on arXiv.
Explaining Reinforcement Learning Decisions in Self-adaptive Systems
No summary available — see the abstract on arXiv.
LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review
No summary available — see the abstract on arXiv.
Wolff-Parkinson-White Detection at 471:1 Class Imbalance: A Leakage-Controlled Study of the Data Bottleneck
No summary available — see the abstract on arXiv.
Metaplasticity as adaptive gradient preconditioning for incremental learning
No summary available — see the abstract on arXiv.
Belayer: Efficient Fault Tolerance for LLM Agentic RL Training
No summary available — see the abstract on arXiv.
Fractional Optimizers Meet Fractal Activation Functions: An Empirical Study of Multi-Scale Optimization in Neural Network
No summary available — see the abstract on arXiv.
Early Cycle Charge Trajectory Generative Prediction and Full Life Cycle Health Management of Iron-Chromium Flow Batteries Based on FlowBD-E1
No summary available — see the abstract on arXiv.
Randomly initialized autoencoders: fixed points and edge-of-chaos
No summary available — see the abstract on arXiv.
Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
No summary available — see the abstract on arXiv.
BDIP-Net: Dual-Interaction Graph Learning for Property Prediction of Bilayer Materials
No summary available — see the abstract on arXiv.
Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions
No summary available — see the abstract on arXiv.
DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance
No summary available — see the abstract on arXiv.
Efficient Neural-Network-Based High-Resolution Radiative Transfer for CO___ Retrieval, and Application to Interferometric Sensing
No summary available — see the abstract on arXiv.
iFuzz-Meta: An Interpretable Fuzzy Learning Framework Bridging Top-Down and Bottom-Up Knowledge Integration
No summary available — see the abstract on arXiv.
SMOPD: Selective Token-Entropy Masking for Dirty-History Multi-Turn On-Policy Self-Distillation
No summary available — see the abstract on arXiv.
Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings
No summary available — see the abstract on arXiv.
Discrete Diffusion Language Models Are Training-Free Multi-Label Classifiers
No summary available — see the abstract on arXiv.
Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade
No summary available — see the abstract on arXiv.
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
No summary available — see the abstract on arXiv.
Do Uncertainty Signals Help? A Systematic Study of Uncertainty-Aware Decoding with Rollback Mechanisms
No summary available — see the abstract on arXiv.
FedImp: Enhancing Federated Learning Convergence with Impurity-Based Weighting
No summary available — see the abstract on arXiv.
Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation
No summary available — see the abstract on arXiv.
P2E-VQ: ECG-linked representation augmentation for PPG via discrete patch retrieval
No summary available — see the abstract on arXiv.
LUNG-KGMM: Knowledge-Guided Multimodal Learning for Lung Cancer Incidence Prediction
No summary available — see the abstract on arXiv.
Vision 38
Deep Analog: Open-Set Film Emulation with Reference-Conditioned 3D LUTs
Reproducing an arbitrary analog film stock's look from a single reference frame is cast as predicting a 3D lookup table (LUT) conditioned on that reference. The standard real-time enhancement recipe — predicting per-image weights over a fixed bank of K LUTs and blending them — is shown to be a gated mixture of experts that inherits gate collapse onto one expert when trained end-to-end for reconstruction; an entropy term, the enhancement analogue of mixture-of-experts load balancing, restores expert utilization and recovers about 1 dB PSNR, though a fixed basis stays closed-set by construction. StyleLUTNet discards the basis and predicts a single LUT as a residual from the reference, self-supervised on procedurally generated color transforms so it generalizes to unseen stocks without paired data or retraining. Wrapped in the Deep Analog pipeline with histogram tone matching and a physics-informed grain and halation renderer, the color stage reaches 22.05 dB PSNR and 0.925 SSIM at 5.2 ms for 1080p and exports a portable .cube LUT; a second failure mode, residual-scale collapse, shares the root cause and yields the general rule that auxiliary regularization must stay subordinate to reconstruction.
On Cross-Validation for Hyperparameter Optimization of Deep Learning Image Classifiers
Hyperparameter optimization (HPO) for deep-learning image classifiers needs a validation signal, and the way that signal is derived is rarely reported. Three protocols — fixed holdout, reshuffled holdout, and 5-fold cross-validation — are compared by absolute performance-estimation error, the gap between the winning configuration's validation AUROC and its test AUROC, with search space, sampler, training procedure, architecture, and test set held identical. Evaluation spans RSNA pneumonia radiographs, binarized HAM10000 skin lesions, and 200-class Tiny ImageNet, using ResNet-18 and ViT-S/16 backbones at several development-set sizes. On both medical datasets every point estimate favored cross-validation, with the largest error reductions at small sample sizes, while on Tiny ImageNet all three protocols showed negligible error and final test AUROC was similar throughout.
Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning
Autoregressive video generators built on diffusion or flow matching tie a rigid training objective to a static sampling schedule, so inference cannot adapt to the sample being produced. Equilibrium Forcing (EqF) drops noise-level conditioning entirely, which decouples learning the denoising field from the sampling procedure and makes training- and inference-time designs modular. Removing the conditioning enables closed-loop inference algorithms that adapt to feedback from the partially generated sample, reportedly surpassing standard noise-conditional denoisers on challenging autoregressive video benchmarks in both quality and consistency, with accompanying analysis of why the data-dependent behavior emerges.
Fast Test-Time Refinement for Robust Learned Image Compression
Learned image compression (LIC) reaches strong rate-distortion performance but inherits the adversarial fragility of deep networks, and test-time refinement has been proposed as a defense without theoretical grounding or evaluation beyond weak threat models. A systematic study finds an asymmetric adversarial trajectory: moving from an adversarial input back toward the benign region is far easier than the reverse, and adversarial examples can often be roughly recovered within only one or two refinement steps, a phenomenon the authors explain with a two-dimensional tube model. They build FTTR, a fast refinement framework, on this observation and argue the robustness comes from contraction of adversarial regions induced by the input-as-label property of LIC rather than from obfuscated gradients, testing against strong adaptive attacks including white-box ones across multiple codecs.
Scale-Consistent Posterior Dynamics for Diffusion Inverse Problems
Posterior sampling for inverse problems with a pretrained diffusion prior needs a conditional score whose intermediate likelihood term is generally intractable. Starting from a family of posterior stochastic differential equations in which one parameter trades deterministic transport against stochastic exploration without altering the marginals, the authors express the likelihood in rescaled clean-image coordinates organized by log signal-to-noise ratio, project diffusion uncertainty through the forward operator to get a noise-conditioned covariance path, and interleave the transport with a frozen-target Langevin corrector; the discretization uses Lie–Trotter splitting with a variance-matched implicit-explicit predictor. They prove marginal invariance, posterior convergence, and a first-order weak error bound, and report competitive super-resolution and deblurring fidelity on FFHQ and ImageNet within 100 score evaluations. A noiseless box-inpainting study shows heavy exploration only reaches its performance plateau when the matched stochastic increment is injected after the stiff likelihood solve.
RigidBench: Evaluating Rigid-Body Physics in Video Generation Models
Video models are increasingly used as predictors of what happens next, but whole-frame similarity scores blend together errors in motion, geometry, object identity, background stability, and appearance, so they say little about whether objects actually move correctly. RigidBench compares each generated continuation against a simulator reference rollout from the same first frame and motion description across five rigid-body tasks, scoring ten separate measurements using per-frame masks, depth, 6-DoF trajectories, and contact data. Evaluating eight models on the same 100 examples, no model leads on all ten measurements, and across model means higher SSIM goes with larger 3D trajectory error (r = 0.89). The benchmark also ships 5,000 training videos with exact simulator state; fine-tuning Wan 2.2 TI2V-5B on them cuts 3D trajectory error roughly 20% with almost no SSIM change, and probes show object position is represented throughout the diffusion transformer and used during denoising.
AdROD: HyperNetwork-based Adversarially Robust Object Detection for Autonomous Driving
Physical adversarial patches can suppress detections from camera-based object detectors, and defenses based on adversarial training or input purification tend to overfit to the attack distributions they were tuned on. AdROD instead generates a fresh detector every frame from low-rank hypernetworks — needing only 1.6% of the parameter footprint of standard hypernetworks — so an attacker cannot obtain the deployed weights in time, and couples stochastic weight updates with distinct input-space transformations for functional diversity. Two serving modes trade robustness against overhead: continuous protection that exploits disagreement between detectors to recover suppressed objects, and an on-demand mode triggered by kinematic discontinuities in tracking. Evaluation on synthetic benchmarks, physically deployed patches, and end-to-end safety tests in the OpenCDA co-simulator shows it beating five baseline defenses while remaining real-time enough to stop the vehicle at an adversarially patched stop sign.
CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration
Medical vision foundation models are typically locked to one specialty and one output type, handling either pathology or radiology and either classification or segmentation. CoM$^3$eT (Co-representation Multidimensional Multitask Medical Transformer) uses attention over multidimensional context to span pathology and radiology, sparse and dense predictions, and two- and higher-dimensional inputs in a single model. It beat other medical foundation models in an open competition across twelve datasets covering classification, segmentation, and report generation, and tuning under 2.5% of parameters matched full fine-tuning, while federated training across hospitals over ordinary internet links and consumer hardware approached pooled-data performance.
LaGSplat: Inferring Physics-Governed Interactive Simulation from Monocular Video Using Latent Lagrangian Gaussian Splatting
LaGSplat (Latent Lagrangian Gaussian Splatting) learns interactive, physics-governed dynamics of a filmed object from one or a few monocular videos, so a user can later push on it with a force that was never measured or seen in training. A low-dimensional latent state serves simultaneously as the generalized coordinate of a learned dissipative Lagrangian and as the conditioning variable of a Gaussian Splatting decoder; because that decoder's primitives are explicit points that move with the object, an image-space force pulls back through the Jacobian into a latent generalized force and enters the equations of motion — something pixel-space convolutional or NeRF decoders cannot support. Tests span rigid to deformable and autonomous to externally forced real systems, with arbitrary forces applied at any time and the response rendered in real time in 2D or 3D, the Euler-Lagrange assumption trading generality for bounded responses where unconstrained predictors diverge.
Self-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation
Adapting a frozen vision foundation model to many visual domains at once is awkward for low-rank adapters, which apply one fixed subspace to every input, and for mixture-of-experts adapters, which add external routers and large expert banks. Self-Routed Tensor Adapters (SRTA) project each input into a low-rank space, derive routing weights from that representation through a learnable domain matrix, and use them to blend slices of a shared Tucker core, producing a per-sample adaptation matrix with no separate gating network; a progressive depth-weighted objective supervises routing across layers. Across five multi-domain classification benchmarks the method matches or slightly beats mixture-of-experts parameter-efficient baselines while using far fewer trainable weights — 2.77M parameters versus 9.52M for MoLoRA in the four-domain setting at rank 64.
UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures
Autonomous vehicles and robots must ship high-dimensional sensor data under tight bandwidth and energy budgets, but the downstream task changes over time, making per-task codecs brittle and per-task retraining impractical. UniTAC is a single learned image codec that spans task-agnostic to task-specialized behavior by abstracting the task into a per-component importance vector — derived, for instance, from gradient attribution of any downstream model — sent as low-overhead side information that conditions both encoder and decoder; the backbone is fixed and the reconstruction stays human-viewable. The paper analyses when a diagonal weighted distortion is task-consistent and builds a Vision Transformer codec whose token-level conditioning implements the weighting natively; at 0.034 bits per pixel on a localized task a single model hits 91.4% accuracy, 1.9 points below a dedicated task-specific codec and well above universal codecs at 76.9%.
CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?
Video world models are meant to sample from the distribution of physical outcomes, yet existing benchmarks score single generations or compare distributions coarsely in learned feature spaces such as FID, never testing whether the spread of outcomes for a specific phenomenon is right. CaliBench scores generations in physically interpretable discrete spaces — a bin index, a die face, a card suit, a colour — for scenes whose reference distribution is known in closed form (Galton boards, Bernoulli forks, dice, cards, lottery, European roulette), and separates scorability (fraction of generations yielding a readable outcome) from calibration (total variation distance from the reference), with a chi-squared significance test. Across nine scenes and six image-to-video models (WAN-2.7, SeeDance-2.0, HappyHorse-1.0, Veo 3.1, Runway Gen-4.5, Cosmos3-Super) at 32 generations each, most scene-model pairs are significantly miscalibrated, concentrating probability on a few outcomes and in the extreme collapsing to one, as Veo 3.1 does on dice, with no model best across all scenes.
PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation
Zero-shot monocular depth models generalize well but blur thin structures and object edges, which the authors trace to the standard pairing of a large-patch Vision Transformer (ViT) encoder with a convolutional decoder — coarse tokenization destroys pixel-level cues that upsampling cannot restore. PXDepth splits the two jobs: a large-patch ViT supplies global scene context while a separate pixel-space predictor built from Context-Modulated Pixel Transformer blocks keeps full-resolution spatial representations throughout prediction. Across zero-shot benchmarks the model reports sharper boundaries and faithful local geometry without losing global depth accuracy, while staying efficient at inference.
The 10th AI City Challenge
A decade-in-review report on the annual competition for intelligent transportation and physical AI, held at ECCV 2026, covering setup, datasets, evaluation protocols, and leaderboard outcomes. The 2026 edition drew 325 registered teams from 26 countries, up from 245 teams and 15 countries the previous year, across six tracks spanning multi-camera 3D perception, traffic safety captioning and visual question answering, anomaly reasoning, text-based person anomaly search, generative traffic video forecasting, and cross-city detection, plus two out-of-domain leaderboards for fisheye violation understanding and pedestrian intent. The common recipe among winning entries combined foundation models with geometric grounding, retrieval or reranking, synthetic data design, domain adaptation, and controlled inference.
Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models
Adversarial prompt tuning cheaply hardens vision-language models, but the authors find robustness on unseen classes decays as training proceeds because the model latches onto pseudo-robust features — shortcuts that do not generalize. ADAPT pairs a target prompt with a pool of decoy prompts that are steered to absorb those shortcut features while the target prompt is constrained to stay orthogonal to them in embedding space, and an accompanying analysis bounds how shifts in pseudo-robust features affect unseen classes. The disentangling prevents robust generalization overfitting, substantially raising adversarial accuracy on classes never seen in training.
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
Video generation is usually scored on appearance fidelity rather than on whether the depicted task actually gets accomplished. Semantic Task Completion Video Generation reframes the problem around the outcome: given a reference image and an instruction, the clip must show the intended result and stay semantically grounded in the reference, with no requirement to depict intermediate steps or preserve conventional appearance consistency. SemComp-Data supplies evaluation instances across six domains, each pairing a reference image, detailed and brief instructions, and an outcome-centric clip, built by a four-stage curation pipeline; SemComp-Bench then has a vision-language model answer structured binary questions and reports an Outcome Achievement and a Generation Reliability score. Representative video generators struggled to reach the intended outcome while keeping task-relevant grounding in the reference image.
Accuracy and Robustness of Model Cascades Under Data Perturbations
Prediction cascades cut energy use by answering easy inputs with a small model and deferring uncertain ones to a large one, which makes the whole design dependent on confidence estimates staying meaningful under real inputs. This study selects an image-classification cascade at the Pareto optimum of accuracy, routing quality, and energy — competitive accuracy with up to a 10-fold decrease in CO2 emissions — then subjects it to static corruptions and sequential perturbations. Three failure modes emerge: corruption breaks the routing signal while the large model would still have helped; corruption degrades both models so deferral no longer recovers accuracy; and under sequential perturbation predictions stabilize but deferral is suppressed, yielding stable but unreliable outputs.
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
Image generation datasets are usually curated one task at a time, which ignores how generative capabilities depend on one another. The described infrastructure pairs three data engines — building relational supervision for text-image grounding, inter-image transformation, and image-knowledge association — with caption experts that align text-to-image and editing supervision, then schedules a multi-stage curriculum that advances task mix, visual concept distribution, data quality, and resolution in the order capabilities are acquired, closing the loop with gap-aware resampling. The pipeline yielded a 440-million-image text-to-image corpus, 120 million editing pairs, and over 27 million image-entity pairs, used to train multimodal diffusion models at 3 billion and 6 billion parameters from scratch and evaluated on CPI-Bench.
20 more specialized papers
- Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders Ze Zhang, Yang Zhang
- Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering Mohammed Abdul Al Arafat Tanzin, Rudzidatul Akmam Dziyauddin
- Zero-Shot Adaptation of Medical Vision Foundation Models for High-Frequency Micro-Ultrasound Prostate Segmentation Ayusha Abbas, Saram Abbas, Kabita Adhikari
- DualMiT-Net: Local-Global Transformer-Convolutional Fusion for Breast Mass Segmentation in Mammographic Regions of Interest Alibek Kamiluly, Milana Muratova, Yash Patel et al.
- Distribution-free false-alarm calibration and chance-corrected spatial evaluation for industrial anomaly detection Jie Deng
- Memory-Bounded Continuation of Greedy Sampling for Continual Anomaly Detection Yoon Gyo Jung, Jaewoo Park, Kuan-Chuan Peng et al.
- Population Structure Analysis of an Inbred Population using Quantitative Shape Phenotyping from Stereo Retinal Photographs Li Tang, Michael D Abramoff
- In Defense of OCTA: The Reconstruction-Utility Gap in OCT-to-OCTA Synthesis Michael Chertok, Alon Tiosano, Orly Gal-Or et al.
- Deep learning-based computed tomography (CT) derived body composition classifier for colorectal cancer patients Eve Harling (James Watt School of Engineering, College of Science & Engineering, University of Glasgow et al.
- CrevasseSeg: A Label-Efficient UAV Crevasse Segmentation Framework Steven Wallace, William D. Harcourt, Richard Hann et al.
- PWLR: Pairwise Witness Local Rejection for Boundary-Aware Out-of-Distribution Detection Chengyao Jia, Ruixuan Wang
- TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening Dong Chen, Kenneth M. C. Cheung
- Convolution-Free Holistic Multivariance Decomposition Layer for Efficient Hyperspectral Image Classification Tensor Networks S\"uha Tuna, \"Ulker Ba\c{s}ar
- Unsupervised Learning of Cell Instances with Generative Routing Pyramids Ziwen Liu, Martin Weigert
- YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition Serdar Yildiz, Abbas Memi\c{s}, Song\"ul Varli
- Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction Yifei Wu, Yicheng Wu, Qiang Ma et al.
- Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia) AlAnoud AllGhayth, AlJawharh AlOtaibi, Jude AlSubaie
- Training with synthetic data for drone detection in thermal imagery Tanel Liiv, Sander Soodla, Nzamba Bignoumba et al.
- Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann et al.
- Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity Alisher Myrgyyassov, Zhen Song, Bruce Xiao Wang et al.
Safety & Alignment 32
IP Protection in the Era of Visual Generative AI: A Survey
Visual generative models create intellectual-property risks spanning unauthorized training, reproduction, extraction, misuse, and redistribution of both data and model assets, and existing surveys organize defenses by lifecycle stage or mechanism in ways that obscure what each method is actually protecting against. The proposed taxonomy instead classifies methods by the risk variable they regulate — Information Exposure Control, Generative Behavior Constraint, and Attribution & Accountability — crossed with a second axis separating data intellectual property from model intellectual property. Protection methods are reviewed under this frame with evaluation protocols realigned to protection objectives, alongside open problems in proactive model-level safeguards, standardized evaluation, robustness to adaptive attacks, and explainable evidence.
MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
Preference training that combines several objectives additively tends to collapse onto whichever objective is cheapest to improve, so a support agent learns to sound warm while offering no real help. MINT changes one step of preference distillation: candidate responses are ranked by their weakest objective rather than a weighted sum, so the best-balanced candidate is distilled over the most lopsided one under an unchanged DPO objective — the p to negative infinity limit of a generalized-mean family. Across cooperative emotional support and adversarial negotiation, min-selection raises both objectives while cutting imbalance; on emotional support the weaker axis rises from 0.37 to 0.64, surpassing human experts and persisting across full multi-turn rollouts. A turn-by-turn analysis shows the correction scales with how imbalanced the reference policy is and lasts exactly as long as that imbalance does.
STAR-FL: Secure Federated Learning with Spatial-Temporal Analysis and Robust Aggregation
Defenses against data poisoning in federated learning struggle both to separate benign from malicious client updates and to blunt poisoned updates that slip through aggregation. STAR-FL attacks both stages: spatial-temporal clustering flags and removes suspicious updates across training rounds, and the learning rate is adjusted during aggregation to limit the damage from anything that evades detection. Experiments on multiple computer vision benchmarks report that the two components work together to cut attack success rates below state-of-the-art defenses against targeted poisoning, with source code released.
Workspace Topology as an Attack Vector in Agentic Coding Assistants
Coding agents with broad filesystem access routinely ingest third-party repositories, creating an opening for instructions hidden inside that code. The authors define "workspace topology" — directory depth, codebase modularity, in-file injection position, and context framing — and measure how each dimension shifts the attack success rate of indirect prompt injection across open-source repositories spanning 10 languages and 6 engineering domains, testing three injection entry points against open-weight models in open-source agent harnesses. Codebase modularity moves attack success rate significantly, with highly modular workspaces proving markedly harder to attack, and context framing plus security cues planted in the workspace also shift the rate. The authors argue this makes an uncontaminated, topology-controlled test environment a prerequisite for trustworthy coding-agent security benchmarks.
Spectral Saliency for Machine Unlearning
Machine unlearning tries to erase the influence of a designated forget-set while keeping the rest of a model's behavior intact, usually by running gradient updates that counteract what was learned. Taking a cue from the Muon optimizer, which normalizes spectral magnitudes to push updates toward rarely used directions, Spectral Saliency Unlearning (SSU) works in the singular-value domain: it thresholds away weak singular components and updates only the directions carrying a confident unlearning signal, with a theoretical argument for the threshold framed as a forgetting-retention trade-off. Experiments span image classifiers, diffusion models, and large language models.
PL-Guard: Probabilistic Logic Reasoning for LLM Guardrails
Guardrails can be framed as policy consistency: decide which policy-relevant facts hold in a prompt-response pair, then decide what those facts imply under the policy. Prompting and LLM-as-a-judge pipelines blur those two jobs together, which yields both unsafe compliance and needless refusals, so PL-Guard separates them — a local model grounds the exchange into predicate probabilities from renormalized True/False token scores, and ProbLog runs probabilistic inference over hand-written symbolic rules. On the XSTest benchmark with an offline Qwen-based evaluator, unsafe compliance drops from 22.0% for the base model to 0.5%, versus 6.0% for an LLM-as-a-judge baseline, though over-refusal rises to 14.4% against that baseline's 5.2%. The intermediate reasoning steps remain explicit and auditable, making the safety-helpfulness tradeoff visible.
Decorrelation Is Not Complementarity: Skill, Not Lineage, Governs Trusted-Monitor Ensembles
Trusted monitoring uses a cheap trusted model to score a stronger untrusted model's actions, and prior work builds diverse monitor ensembles by minimizing average pairwise correlation — but those twelve monitors shared a base model, leaving the source of diversity untested. Studying 24 open-weight monitors across nine pretraining lineages and a 29x range of detection skill on backdoored code, the authors decompose agreement on attack items into a shared-detectability signal component and an idiosyncratic error component that predict ensemble gain with opposite signs (Spearman -0.25 and +0.26), so the correlation metric actually used for panel construction predicts gain barely at all (+0.05). Pretraining lineage does not buy usable decorrelation either: at matched member capability, cross-lineage panels detect no better, and panel gain over the best single member falls monotonically with panel skill, with no correlation-weighted selection beating simply picking the strongest monitor out of sample. The authors report an earlier version of their own analysis that flipped when two monitors were added, arguing the effect is a property of the assembled pool.
A Privacy Study of Sparse Collaborative Inference
Collaborative inference splits a neural network between an edge device and a server, and recent systems sparsify the transmitted intermediate activations to cut bandwidth — a step often assumed to also improve privacy. The authors test that assumption by separating a sparse activation into its retained values and the set of positions those values occupy, then attempting input reconstruction from each part alone. Sparsification reduces privacy leakage far less than it reduces transmission cost, and the positions alone — usually treated as mere decoding side information — support high-fidelity reconstruction and re-identification of individuals across natural-image and face datasets, even when both bandwidth and task utility are low.
SAUL: Sharpness-Aware Augmented-Lagrangian Unlearning
Unlearning targeted knowledge from a large language model usually degrades its general capabilities, because the desired amount of forgetting is only expressed implicitly through the loss. SAUL (Sharpness-Aware Augmented-Lagrangian Unlearning) instead states forgetting as an explicit constraint with a pass/fail criterion, using an augmented Lagrangian controller that scales forget-side pressure to constraint violation and switches the forget update off once the criterion holds, plus sharpness-aware updates and separate optimizers for the retain and forget objectives. On TOFU, WMDP, and MUSE it improves the forgetting-utility trade-off over sharpness- and perturbation-based baselines, and the Lagrangian controller alone, dropped into existing methods, raises their post-forgetting utility.
Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors
Machine unlearning removes a subset of training data from a model to satisfy privacy legislation, and existing work concentrates on hand-designing an unlearning function whose output approximates a model retrained from scratch — an approach whose intricate structure becomes the bottleneck on large datasets even for small models. Borrowing from Learning to Optimize, L2UL instead learns the unlearning behavior itself from a distributional view, yielding a simple, model-agnostic unlearning function obtained by training rather than by design. Experiments report accuracy comparable to full retraining with substantially better efficiency in data-intensive settings, with scalability checks on larger ResNet models.
The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human Feedback
The Ethical Decision Head encodes normative ethical reasoning as a differentiable reward for a driving policy trained with Proximal Policy Optimization on states aligned with the CARLA simulator, using a Bradley-Terry reward model fit to pairwise human preference annotations over 200 collision-imminent scenarios. Two frameworks are instantiated: a Utilitarian one minimizing total casualties and a Kantian one treating course maintenance as a categorical imperative — the latter collapsing to a constant prediction that serves as a pipeline control confirming training stability. The utilitarian agent revealed an asymmetry in what human supervision can teach: raters rewarded vehicle self-sacrifice over casualty minimization, and the policy learned that preference faithfully, suggesting reinforcement learning from human feedback captures ethics as people actually reward it rather than as they profess it.
Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution
Once an agent can write files, send messages, and launch jobs, the safety problem moves from harmful text to harmful side effects, and prompt-level policy shapes behavior without creating an actual execution boundary. Aegis treats model outputs as action proposals and interposes a trusted decision layer before any tool runs: proposals are checked against active policy state, provenance is resolved server-side, uncertainty fails closed, and selected cases route through a quorum-based, non-unilateral authorization path. Over a sandbox corpus of 6,300 rows, prompt-policy conditioning alone produced 79 risky leakage rows, whereas the 2,100 Aegis-governed rows recorded zero governed mock-tool applications and zero risky side-effect completions, with all 1,832 attempted governed rows preserving trusted provenance — a systems claim about this corpus, not a general autonomy-safety guarantee.
Position: Fairness Failure in Generative Models is an Evaluation Problem
Concerns that generative models reinforce societal inequalities have persisted without becoming actionable, and the argument advanced here is that the root cause is measurement: fairness results are rarely comparable across papers or usable for deployment decisions, so the failure is an evaluation problem rather than only a modeling one. Recurring empirical and conceptual failure modes in current practice are catalogued to motivate replacing ad-hoc bias checks with standardized, generative-specific evaluation. The concrete proposal is Fairness Cards, a minimal reporting artifact that records prompt families, counterfactual protocols, metrics, and refusal handling so results become reproducible, comparable, and attributable.
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
Defenses against Not-Safe-For-Work generation in text-to-image models mostly require white-box access — editing text encoders, weights, or inference internals — which rules them out for proprietary APIs, while black-box prompt rewriting breaks down on what the authors call benign adversarial prompts: wording that reads as safe yet still triggers harmful output because of what the model learned. DiSCO is a zero-shot plug-in operating purely on prompts, appending suffixes found by beam search scored contrastively against pools of safe and unsafe images the target model itself generated, iterating with adaptive feedback until output is clean. On the I2P benchmark under several red-teaming attacks it cut attack success rate by 37.7% on undefended models and 25.13% on already-defended ones, while preserving semantic fidelity and improving image coherence.
Authorization Before Context: A Model-Neutral Audience Boundary Against Cross-Audience Memory Leakage in Agentic Systems
A personal language agent that learns a fact from one audience can later place it in a prompt assembled for a different one, an attack surface opened by ambiguous channels, cross-audience prying, and poisoned memory. The proposed rule, authorization before context, tags each stored item with the audience present when it was recorded and admits it only when every current viewer already belonged to that audience, defaulting to public when channel metadata is ambiguous. The authors prove this anti-monotone rule preserves cross-channel recall while making leakage structurally impossible rather than dependent on model behavior, and report that no forbidden fact entered the context their boundary assembled on a synthetic Contextual-Integrity suite, with all read paths audited to fail closed. The evidence is preliminary and synthetic.
Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents
Retrieval-Augmented Generation (RAG) systems can be steered by poisoned retrieved documents, and a model will often correctly flag a document as false while still being influenced by it. Instead of the Cordon Principle's costly isolation of the answer-synthesizing model from raw evidence, the refined principle here is that only agents capable of deliberative System 2 reasoning should be allowed to access untrusted documents. New metrics quantify the gap between detecting misinformation and being swayed by it, and comparisons across current models find reasoning-capable ones substantially more robust to corrupted evidence without needing strict isolation.
The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence
An AI audit record is only useful if its durability and trust boundary are stated explicitly, since returning a guarded decision before any durable write minimizes latency but cannot guarantee the evidence survives an immediate crash. The rebuilt RuntimeGuard-AI prototype binds each deterministic policy decision to the exact policy source, commits a privacy-minimizing record at a caller-selected synchronization boundary, and returns an Ed25519-signed receipt stating whether that boundary completed, with committed records grouped into chained, signed Merkle epochs an auditor verifies using an externally obtained key. On an Apple M4 Pro with four worker threads and 2,048-byte prompts, buffered signed evidence reaches 27,193 requests per second at 141.9 microseconds median latency, while full per-record synchronization drops throughput to about 242 requests/s and raises median latency to 16.0 ms. The authors frame this as a measured durability-latency trade-off and note the prototype does not prove model execution, stop a compromised signer from forking history, or establish legal conformity.
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models
Safety, security, and compliance benchmarks were designed for large models, and it is unclear whether their automated pipelines transfer to the small language models now deployed in resource-constrained, privacy-sensitive settings. Running five widely used benchmark suites against 26 open-source small models under a single rubric that scores each response as harmful, safe, or ambiguous shows ambiguous judgments dominate, increasing with lexical density, output perplexity, and output length, and decreasing with lexical sophistication, self-coherence, and reply-prompt similarity — a confound that mixes model capability with apparent safety. Because ambiguity is so prevalent, mean-score leaderboards are brittle: model rankings shift substantially under different but equally reasonable treatments of ambiguous responses, even when the underlying outputs are identical.
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models
Safety alignment in open-weight language models can be stripped in minutes by abliteration, which projects the refusal-mediating direction out of the weights, and no known release-time defense prevents this durably. Decoy hardening (Fool's Gold) concedes the strip and poisons its payoff instead: trained inside a differentiable simulation of the attack, the model answers hazardous operational requests in the attacked state with confident, fluent responses whose critical elements are falsified, while a refusal pin and benign leash hold clean-state behavior close to the original. Across seven models from five families (9B-122B, dense and mixture-of-experts), six passed a pre-registered efficacy gate with 0.51-0.90 of attacked-state responses being decoys; on a chemical and biological red-team slice the defended 122B model is fatally wrong on 0.82-0.86 of matched-quality answers versus at most 0.10 undefended, and element-wise consensus over 64 samples reconstructs a usable procedure far less often than undefended, with no label-free way to distinguish the regimes. The defense covers only the initially released weights and does not address in-context jailbreaks.
PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance
Agents that plan decentralized finance (DeFi) transactions inherit their language model's susceptibility to prompt injection, and nothing normally binds a verifier's approval to the exact bytes submitted on-chain. PACE interposes a transaction-level authorization layer with typed transaction intents, a deterministic policy verifier, and signed Policy Decision Records that cryptographically tie the approved intent, policy, and simulation report to the execution payload, with replay and expiration protection enforced by a Solidity smart account at roughly 30,000 gas of overhead. Across 40 tasks and 2,800 trials in a deterministic sandbox, it reports a 0.00 unsafe-execution rate against 0.80 for the unguarded baseline, with no false positives on benign tasks; ablations single out permissive policy settings and the touched-contract allowlist as the dominant safety components. The authors frame the claim as logic-level safety within a reproducible benchmark rather than deployment-ready DeFi security.
COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models
In many multimodal jailbreaks neither the prompt nor the image is harmful alone; unsafe behavior appears only when the model binds an innocuous operation such as summarizing or translating to a specific visual target, which defenses that moderate the prompt-image pair as a whole tend to miss. COMIC is a pre-generation safety gate that infers the requested operation and reference type, assembles candidate targets from optical character recognition and open-vocabulary proposals, grounds the plausible referents, and evaluates safety over explicit operation-target pairs, combining max-risk aggregation with quality-aware routing to stay conservative when grounding is ambiguous. Across several open-source multimodal large language models and both localized and broader jailbreak benchmarks, it improves robustness while preserving utility on benign reference-sensitive inputs, supporting the argument that multimodal safety requires modeling the operation, its visual target, and the confidence of that grounding.
Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets
Jailbreak papers report attack success rate without pinning down how many attempts it took, so methods that differ tenfold in budget get compared head to head, and compute-normalized alternatives based on FLOPs are unmeasurable for black-box targets. Fair-ASR fixes the comparison axis at the number of calls to the target model while tracking attacker-model calls separately; re-running 11 representative attacks under it shows rankings shift substantially as the target-call budget changes, with simple stochastic perturbations and hand-written templates staying competitive and no language-model-driven method being efficient in both call types. The authors exploit that gap with ReCode, which composes desensitization rewriting with two cheap primitives and reaches 85% attack success on GPT-5 within 20 target calls at 7.19 attacker calls per request.
Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services
Stateless safety filters judge one request at a time, so an attacker can split a harmful task into individually permissible requests; a stateful monitor could catch this only by grouping those requests, which fails once the attacker uses unlinkable identities and recombines the answers elsewhere. A theoretical analysis shows that for a fixed attack strategy without retries, the achievable security-utility tradeoff depends entirely on whether benign requests for the same capabilities form persistent, recognizable groups — and that once attackers can retry and learn from Allow/Block decisions, even that useful operating point disappears, since the feedback reveals what passes but not whether a block was correct. Experiments on 91 executable tasks and 11,393 capability-matched benign requests bear this out: under a 1% denial cap on those requests and 0.5% on background traffic, all ten tested policies, including a privileged one holding an exact request-to-operation map, either failed to stop the attacks or exceeded the budget, with attack success at least 99% after one attempt on unseen task families. Workable defenses, the authors argue, need reliable identity linkage, costs for minting fresh identities, or control over how answers are used.
Effects of Answer Format Variation on Gender Bias in Large Language Models
Survey science has long known that answer format shapes responses, and language models are similarly sensitive to prompt wording, yet bias benchmarks generally fix a single response format. Three instruction-tuned models are evaluated on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled, and open-ended formats under otherwise identical conditions, comparing both measured bias and alignment with human response distributions. Format substantially altered the measured outcomes, including reversals in how models rank against one another, because each format elicits a different behaviour — forced choice, scale-based distributions, or refusal in free text — which argues for treating format as a substantive part of evaluation design.
Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings
Prompt-safety guardrails built on LLM-as-a-judge or cloud moderation APIs add roughly 250-900 ms per request and route user text through external endpoints, which rules them out for systems that must answer in under 100 ms. Reflex-Guard runs locally, combining jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers. On a balanced 30,568-sample set drawn from five sources it reaches 95.9% recall on harmful prompts at 37.6 ms end-to-end latency, against 255 ms for Llama Guard 2 and 723 ms for SafeDecoding, and detects all GCG suffix attacks and Base64-encoded prompts at the default threshold. DrAttack-style structured prompts occupy a distinct region of the embedding probability space and needed the threshold lowered to 0.03.
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
Agent harnesses manage tools, extensions, persistent state, permissions, and external actions, but existing safety benchmarks target individual attack mechanisms rather than these operational responsibilities, making failures hard to compare. HarnessRisk organizes 128 sandboxed cases across six lifecycle phases — harness configuration, capability extension, runtime operation, state persistence, action control, and incident recovery — each pairing a benign user objective with an adversarial instruction embedded in an untrusted workflow artifact, scored on utility, attack success rate, persistence, and detection. Across three harnesses, six language models, and 14 model-and-harness configurations, attack success ranges from 12.6% to 80.9% while utility stays between 75.0% and 97.6%, harness configuration is the most vulnerable phase in every harness, and explicit risk recognition does not translate into safe action: some configurations flag risk in over 90% of runs yet still succumb.
MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps
Graphical user interface (GUI) agents that drive real smartphones must read untrusted on-screen content, which exposes them to environmental injection — hidden instructions arriving through notifications, app content, and other everyday channels. MobileWorldSafety is a benchmark of 142 risk tasks on real Android apps, each with a programmatically verifiable risk indicator over the final system state, scored by rule-based checks with an LLM judge for ambiguous cases so that safety failures are separated from plain capability failures. Across six general-purpose and GUI-specialized agents, attack success rates ranged from 40.4% to 66.9%, indicating none reliably hold their safety alignment when adversarial text is presented as ordinary mobile context.
GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities
Communities of LLM agents that debate on social platforms can be pushed toward group polarization, but existing attacks require editing agent prompts or constructing echo chambers, neither of which is easy against a live deployment. The proposed Memory-Mediated Polarization Cascade instead exposes a few target agents to stance-reinforcing arguments that their memory systems retain, then uses a stance-neutral public discussion to cue retrieval so those arguments resurface and spread to untreated agents; GraphWake implements this with stance-support argumentation knowledge graphs, axiom-oriented triple selection, and neutral memory cueing. Across multiple discussions and memory systems the attack substantially increases group polarization, revealing a community-level risk that runs through agent memory rather than prompt access.
An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning
Machine unlearning for language models is normally scored on two axes — suppressing target knowledge and preserving unrelated utility — which leaves undefined what a model should do with a prompt adjacent to the forgotten topic that could be answered generally without leaking specifics. Using a controlled LoRA-GRPO setup on the RWKU benchmark, the study compares four reward designs spanning lexical suppression, anti-refusal shaping, rubric-based broad answering, and explicit refusal contrast, with and without supervised fine-tuning warm-up. Optimization success turns out not to imply behavioral unlearning: forget scores, held-out completion audits, terminal rollout audits, and training dynamics frequently disagree, which the authors trace to reward hacking, policy-support limits in GRPO, and benchmark probes that miss changes at the endpoints they measure.
The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges
System prompts and retrieved documents give models useful context but create an attack surface, since adversarial inputs can coax the model into disclosing them; prior detection work reads hidden states, which is awkward to deploy. LeakGauge instead appends a suffix that probes leakage behavior and maps the resulting prefill token probabilities to an attack-risk score, and finds that a content-agnostic gauge verbalizing leakage behavior beats one built from the confidential text itself. Across 11 models including GLM-5.2 and Kimi-K3, it reaches AUROC between 0.944 and 0.996 on unseen attacks, holds up when content changes language or attacks shift from verbatim to semantic disclosure, and supports an input detector adding under 0.5K parameters and 10.34 ms latency; activation steering along an internal leakage direction moves the score, tying the observable signal to internal representations.
2 more specialized papers
- When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice Muhammad Salar Khan, Hamza Umer, Hasan Mahmud et al.
- Traceable Trust for action-ready artificial intelligence in bioscience Huayu Xin, Yizhi Cai, Mukilan Deivarajan Suresh et al.
Reinforcement Learning 25
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
Chess platforms serve millions of automatically generated tactics puzzles, but unlike expert-curated sets they are widely considered to have little teaching value, and recommendation relies on heuristics. Using 1.5 billion puzzle-solving histories from a full year of user data, the authors learn a puzzle's pedagogical value and an offline reinforcement learning policy for selecting practice sets that balance tactical motifs and required look-ahead depth. Offline policy evaluation indicates significant gains for beginners in the 100-1000 puzzle Elo range, especially those whose learning had stalled, and expert chess players were recruited to rate the puzzles the model surfaces. The authors frame the pipeline as a general route to inferring the instructional value of practice items from ordinary user interaction logs.
Do Geometry-Aware Positional Encodings Help Transformers in Spatial Imperfect-Information Games?
Transformers playing spatial imperfect-information games must encode map geometry while tracking entities they cannot see, raising the question of whether geometry-aware positional encodings actually help. A four-level benchmark on a hexagonal naval pursuit game — geometry and topology probes, an exact-Bayes hidden-target tracking task, offline imitation at 1k and 10k games, and 7,200 fixed-seed matches against three legacy opponents — compares HexRoPE, rectangular relative bias, graph bias and no encoding on matched backbones. HexRoPE cuts exact-belief posterior cross-entropy by 0.278 to 0.329 with confidence intervals excluding zero, and improves imitation action accuracy by 4.63 points at 1k games shrinking to 1.55 points at 10k, but its effect on aggregate win rate is -1.56 percentage points with a confidence interval spanning zero, so better belief representation did not translate into stronger closed-loop play. Rectangular relative bias led on symmetry-consistent belief yet failed sharply when extrapolating from radius 3 to radius 4.
PureTD: Reinforcement Learning for Backgammon Money Games with No Evaluation-time Search
Tesauro's TD-Gammon showed self-play temporal-difference learning could master backgammon, but strong modern engines still rely on look-ahead search when choosing moves. PureTD learns both checker play and doubling-cube decisions from scratch through self-play reinforcement learning, with no expert features, minimal hand-coded logic, and no search at evaluation time. In cubeful money games the search-free model is both faster to evaluate and substantially stronger than GNU Backgammon and Open Sage running one-ply look-ahead, reaching near-state-of-the-art playing strength.
Temporal Logic Guided Universal Task Representations for Reinforcement Learning
Task-conditioned reinforcement learning agents usually encode tasks with bespoke representations that do not transfer across settings and that are trained only through the controller's gradients, which limits representation quality. LOTUS encodes tasks written as linear temporal logic (LTL) formulas with an architecture built to capture the relations among subformulas, and treats the LTL encoder as a policy in its own right rather than a passive module updated by controller gradients, with a bisimulation metric supplying guarantees on behavioral equivalence, optimality fidelity, and trajectory robustness. It drops into any RL algorithm and reports success-rate gains of 15–45% on unseen manipulation tasks, over 20% faster convergence on single tasks, and more than 25% better generalization on multi-task settings with deeper subgoal chains or more conjunctions.
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Reinforcement learning with group-relative advantages is standard for post-training language model reasoners, but when several reward objectives are involved the usual recipe collapses them into a fixed weighted sum before group-wise standardization. That design lets rollouts with very different reward profiles receive identical advantages and keeps spending gradient budget on objectives that are already saturated. SA-MRPO standardizes each objective separately and discounts its contribution by a batch-level saturation estimate, shifting effort toward objectives with remaining headroom — and the authors show this can flip an update's sign, not just rescale it. Against GDPO, the method improves the harder correctness objective in 12 of 15 benchmark comparisons, with gains up to 5% on AIME24, 9.2% on AMC23 in adaptive reasoning, and 2.3% on coding pass rate, while holding the easier objectives near their satisfied levels.
Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics
Deep Q-learning trains unstably, and existing explanations isolate single causes such as overestimation bias without describing how they interact during recursive value estimation. The analysis here splits instability into three levels: operator-level bias in Bellman bootstrapping, estimator-level sensitivity of greedy action selection to regression noise, and parameter-dynamics imbalance under aggressive data reuse, identifying a reward-triggered self-reinforcing trap and characteristic parameter spike dynamics. From these, the authors derive stabilization principles — controlled bootstrapping, ensemble quantile estimation, and spike-based parameter regulation — and report competitive performance with steadier training on Atari-100K and Procgen.
SCALE: State-Calibrated Latent Embeddings for JEPA Planning in the Right Geometry
World models that plan by comparing predicted terminal embeddings to a goal embedding depend on the geometry of the representation, and the two dominant recipes — inheriting a pretrained feature space as in DINO-WM, or learning end-to-end with anti-collapse regularization as in LeWM — succeed on different tasks. The authors observe that although task state is decodable from both models' full embeddings, DINO-WM's leading principal components carry much more state information, which matters because Euclidean planning costs are dominated by high-variance directions. SCALE (State-CAlibrated Latent Embeddings) adds one training-time regularizer correlating sampled pairwise latent distances with distances in a standardized task-state space, and improves every task-solver average over LeWM across five tasks, three planning solvers, and five compute budgets with no planning-time overhead — while a latent-to-state regression control that matches decodability but not distance alignment gives inconsistent gains.
Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation
Fluid and vasopressor dosing in sepsis is a sequential decision under uncertainty, but a learned policy cannot be tested on patients, so everything hinges on off-policy value estimates that are notoriously fragile and optimistic. Using 36,872 septic ICU stays from MIMIC-IV, dosing is cast as a discretized Markov decision process with 1,000 states and a 5×5 grid of 25 fluid-vasopressor actions solved by policy iteration, with the clinician behavior policy fit by a random forest and evaluation done twice, via weighted importance sampling and fitted Q evaluation, alongside reliability diagnostics. Modeling the behavior policy with a random forest lifted the effective sample size from 4.0 to 50.1, and both estimators place the learned policy above observed clinician return (50.8 and 46.8 versus 38.2) while departing only modestly from practice, mainly by giving less intravenous fluid.
Le Critique: Privileged Value Functions for LLM Reinforcement Learning
Reinforcement learning for large language models mostly differs in how it reduces gradient variance: group-relative methods like GRPO sample many rollouts per prompt but give only sequence-level credit and stall on straggler rollouts, while learned value functions would give token-level advantages without large groups but have been hard to justify given their infrastructure cost. Two complementary fixes are proposed — Privileged Value Functions, which feed extra task-relevant token-level signal into the critic without biasing the policy objective, and TETHER, a baseline that adaptively interpolates between the group-relative and value baselines according to how accurate the value function currently is. Across several reasoning tasks both consistently improve on the standard value-function baseline and are competitive with or better than mean-baseline GRPO.
Q-based Variational Inverse Reinforcement Learning
Hand-specifying human preferences for an AI system is usually infeasible, so inverse reinforcement learning (IRL) infers reward functions from expert behaviour instead — but Bayesian variants that quantify uncertainty have not scaled. Q-based Variational IRL (QVIRL) recovers a posterior over rewards by primarily learning a variational distribution over optimal Q-values, combining scalability with the uncertainty estimates that safety-critical deployment and active learning need. Apprenticeship learning results span gridworlds, Lunar Lander, the Highway Environment, and two ATARI games under both static expert data and active querying, making it the first Bayesian IRL method demonstrated to train from raw pixel observations.
Q-Learning With World Models
Model-based reinforcement learning typically trains policies or value functions on imagined rollouts, which compounds model bias and scales poorly to long-horizon, visually complex problems such as real-world robot manipulation. QWM keeps policy and value learning grounded in real online transitions and uses the world model only for test-time search over imagined trajectories, selecting high-value actions during both rollouts and evaluation. On the Robomimic and LIBERO manipulation benchmarks it outperforms strong prior state-of-the-art methods on both sample efficiency and final performance.
Task Specialization Fine-Tuning for Contextual Reinforcement Learning
Contextual reinforcement learning aims to cover a whole space of related tasks, and the paradigm proposed here pretrains a single policy and then fine-tunes multiple specialized ones, which raises the question of how to split a constrained budget across task regions with unequal marginal returns. Task Specialization Fine-Tuning (TSFT) predicts each region's fine-tuning payoff with a simple parametric model and solves the resulting discrete budget allocation exactly via integer linear programming, online. Across combinatorial optimization, continuous control, and LLM fine-tuning, it significantly outperforms baselines on task coverage and approaches oracle allocation.
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Reinforcement learning for reasoning still depends on ground-truth supervision such as verifiable rewards, and self-rewarding alternatives tend to reinforce a model's existing biases, shrink response diversity, and collapse during training. Co-RL trains several decoupled models sharing no parameters at once, each optimized against rewards derived from its peers, and shows that raising cohort diversity through heterogeneous model families, sizes, and rephrased training samples reduces the correlated errors that drive self-reinforcing feedback loops. Using no ground-truth labels, it reports average gains of 3.0-8.6% across seven text benchmarks for language models and 2.3-7.2% across four multimodal benchmarks for vision-language models, matching or surpassing supervised methods while preserving behavioral diversity.
Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning
Image-based reinforcement learning burns samples on redundant updates because replay selection is random or poorly targeted, and prioritized replay and intrinsic exploration rewards have largely been studied apart. NSPER scores stored transitions by novelty, which surfaces underrepresented states, and by surprise, which flags where the agent's model of the environment is wrong; the NSPER+R variant feeds the same two signals back as intrinsic rewards so replay quality and exploration improve together. On DeepMind Control Suite tasks both variants converge faster than existing prioritization methods.
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
Training coding agents with reinforcement learning through their production harnesses breaks in specific ways: sandboxes crash, agents hack rewards, and the tokens a harness actually sent diverge from what the trainer scores. LEGO-RL bridges the two without touching harness control flow, proxying LLM calls in-process to capture raw generation streams for token-level alignment and trainer-side log-probability recomputation, orchestrating cached sandboxes with stage-wise defenses against reward hacking, and adding automated validation plus a live trajectory UI. Training the sparse mixture-of-experts model Qwen3.5-35B-A3B with GSPO lifted SWE-bench Verified scores across three harnesses — OpenHands SDK from 64.0% to 70.4%, Claude Code from 62.4% to 68.2%, and OpenCode from 57.2% to 66.6% — while keeping rollout-versus-training probability correlation above 0.99.
Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context
Reinforcement learning for tool-using conversational agents usually collapses an entire rollout into a single terminal reward, giving identical credit to effective elicitation, to errors, and to the repairs that follow. Feedback-Aware Credit Assignment (FACA) treats each user reply as noisy, temporally local evidence about the preceding user-to-user segment, converts it into a locally normalized reaction advantage, and adds it to the verified terminal outcome advantage without an extra critic or additional rollouts. Against an outcome-only Interactive GRPO control matched on simulator, visible dialogue, initialization, rollout, and optimization, it raised the nine-domain tau-family average by 5.91 points at 8B and 10.22 points at 14B across three independently trained runs, with the same ordering holding zero-shot on Pare-Bench and Co-Gym; randomizing the polarity of user reactions removed the largest gain, indicating the signal comes from the feedback content itself.
Agent Lightning v1.0: Towards Harnessed Agentic RL
When reinforcement learning trains a model through an existing agent harness, the harness rather than the training engine owns the environment interaction loop, and the trainer observes only sequences of LLM request-response pairs — which creates problems in retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling that affect training stability. Agent Lightning v1.0 implements this harnessed agentic reinforcement learning pattern in roughly 3,500 lines of code, connecting arbitrary harnesses to training via an LLM endpoint proxy, an architecture since adopted by verl Uni-Agent, AReaL 2.0, slime, and Polar. Evaluated on instruction-following, search, and coding agents, it took Qwen3.5-9B from 41.8% to 56.4% on SWE-bench Verified using only 6K training examples and modest compute, with the complete workflow and training scripts released for reproduction.
No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models
Joint-Embedding Predictive Architectures (JEPAs) can collapse to a constant encoder, and systems such as LeWM block this with SIGReg, a regularizer that forces the latent distribution to match an isotropic Gaussian regardless of the environment being modeled. AC-MTM (Action-Contrastive Masked Transition Modeling) instead derives the anti-collapse pressure from the transition data: a training-only inverse-dynamics head must identify which action produced each latent transition among the batch's other actions, something a collapsed encoder provably cannot do, and the head is discarded afterward so test-time encoding, planning, and compute match LeWM exactly. Across four pixel-control tasks it trains stably from scratch and matches SIGReg on average, while on the multi-object OGBench Visual Scene task it reaches 80.0% success versus 58.0% for SIGReg, consistent with the prescribed Gaussian geometry becoming a bottleneck. The signal needs no target network, stop-gradient, pretrained encoder, or reconstruction loss.
Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation
Reinforcement learning with verifiable rewards (RLVR) gives easy and hard training samples the same rollout budget, and existing adaptive schedulers need difficulty estimates that either cost extra probing generations or suffer cold starts, stale feedback, and blindness to relations among samples. The proposed plug-and-play estimator builds a sample graph from semantic and reasoning similarity, assigns latent difficulty states under a Potts prior so neighbors tend to share state, aggregates rollout outcomes per state with a Beta-Binomial model, and updates assignments online through mean-field variational inference. Sharing rollout feedback across related samples supplies continuously refreshed difficulty estimates without any dedicated probing, and dropping it into both sample-selection and rollout-allocation schedulers improves results across several base models and benchmarks.
Towards Zero-Shot Task Transfer with Neurosymbolic World Models
Model-based reinforcement learning usually plans in an uninterpretable latent space whose representations are entangled with the training task, making transfer hard. The proposed world model restricts reward prediction to a structured, symbolic subset of the latent state, decoupling observation reconstruction from reward prediction. The resulting models adapt zero-shot to new reward functions defined over the same symbolic state space, without any further environment interaction, and generalize considerably better than purely neural world models, though the paper also lays out the added difficulty of learning such neurosymbolic representations.
Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents
Systems that pair a large language model planner with a reinforcement learning controller are increasingly common, but the theoretical standing of LLM-derived reward signals is usually left implicit. Recasting the architecture as a Goal-Augmented Markov Decision Process and treating the LLM's per-state progress score as a bounded potential function shows that the resulting shaping term preserves the set of optimal policies even when the LLM's scores are inaccurate, a stronger guarantee than generic LLM-as-reward approaches provide. Numerical verification on a small MDP covers four potential configurations, including an adversarial one scaled to twenty times the base reward magnitude.
4 more specialized papers
- Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning Joanikij Chulev, Hendrik Baier
- Continuous Quantum Feedback Control via Kraus-Parameterized Belief Reinforcement Learning Priyanshi Singh, Krishna Bhatia
- Learning Stock Trading Policies via Barycenter-Based Adversarial Inverse Reinforcement Learning Arishi Orra, Himanshu Choudhary, Manoj Thakur
- Self-Supervised Auxiliary Task Discovery for Stable Reinforcement Learning in Stock Trading Arishi Orra, Himanshu Choudhary, Manoj Thakur
Multimodal 14
Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift
Audio-text foundation models collapse under severe acoustic noise, and existing fixes either adapt by gradient descent at test time, which reinforces the noise, or require noise annotations that are unavailable at inference. PRISM is a training-free, source-free adaptation method built on the observation that heavy noise induces a low-rank affine shift in the shared audio-text latent space, with over 90% of the distortion energy in the leading 60 principal components; it estimates and reverses that shift from an unlabeled target batch using frozen text prototypes as geometric anchors, compiling three closed-form corrections into one static projection matrix. On UrbanSound8K it beats the zero-shot baseline by 12.94 percentage points and an oracle-assisted test-time adaptation baseline by 9.41 points, while costing a single matrix-vector multiply at inference. The authors also identify a failure mode for broadband polyphonic classes and recover up to 8.16 points on the worst-affected class with a confidence-aware variant.
Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer
Text-to-audio-video models generate a clip and its soundtrack from a prompt but give no control over who is speaking. Adding a single zero-initialized linear layer on top of the audio backbone, fine-tuning briefly, and conditioning on a short reference recording turns such a model into a voice cloner; the reference enters through two paths, its diffusion latents prepended to the audio stream and a global speaker embedding modulating the target audio tokens. On 674 speaker-text pairs covering 30 speakers, the enhanced 5B model achieves the highest speaker-encoder cosine similarity under ECAPA-TDNN, WavLM-SV, and Resemblyzer, significantly beating all five text-to-speech cloning baselines. Because the audio path can run without the video path at inference, generation is roughly 30x faster than the full audio-video diffusion loop while keeping the cloned voice.
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
Visual chain-of-thought (Visual CoT) lets multimodal models reason by generating intermediate images, but synthesizing and re-encoding those frames is expensive, which hurts proactive video reasoning where the model must anticipate what happens next. Internalized Visual Thinking (IVT) is a post-training recipe that jointly optimizes text prediction and next-embedding prediction over unlabeled video, so the model learns to predict latent representations of future frames — motion, object transitions, interactions, intent — during training but answers directly at inference with no image generation. Across controlled studies of target representations, decoder designs, prediction horizons, data mixtures, curricula, and objectives, IVT beats direct-answer fine-tuning on all six evaluation settings and matches or exceeds explicit Visual CoT while cutting average end-to-end latency by more than 5x, suggesting pixel-space generation at inference is not required for this kind of reasoning.
The Limits of Binding in Dual Encoders
Dual encoders like CLIP score an image-caption pair with one inner product between independently computed vectors, and score near chance when asked to tell a red car with a blue dog from a blue car with a red dog. Working inside an existing ideal-encoder axiom framework, the authors first show those axioms are satisfiable — so any impossibility must come from an extra, checkable hypothesis — then prove three obstructions: recursive role-binding codes obey an exact swap-margin law m(D) = 2b^-D in nesting depth, making resolvable depth grow only logarithmically in dimension and land in single digits at CLIP scale; the contrastive objective's entire reward for binding is capped by how often training contrasts a caption against its own swap, a rate that vanishes at web scale; and a tight frontier trades binding margin against how closely swap-related captions must sit to a shared paraphrase anchor. A text-only diagnostic across 18 deployed encoders puts every model at roughly 25-35% of its own ceiling, with the induced per-item ceiling tracking SugarCrepe subset difficulty at r = 0.99 — meaning today's failures are an incentive and code-structure problem, not a dimension or smoothness limit.
Uncertainty-Aware Decision Making in Multimodal Large Language Models
A survey of uncertainty in multimodal large language models (MLLMs), organized not around confidence scores but around what a system should do with them. The framework runs from uncertainty sources — bad inputs, perceptual errors, weak grounding, cross-modal conflict, distribution shift, unanswerable questions — to observable signals, to calibration and risk control, to concrete actions, covering token and logit uncertainty, semantic disagreement, perturbation instability, grounding and attribution scores, verbalized confidence, verifier and judge scores, conformal prediction, selective answering, abstention, clarification, retrieval, self-checking, and escalation. The central argument is that uncertainty methods should be judged by whether they change system behavior under insufficient or conflicting evidence, not by calibration numbers alone, with open problems named in source-aware decomposition, action-aware benchmarks, calibration under shift, and black-box estimation.
Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
Whether multimodal models recognize a sad voice and a sad face through the same internal machinery or through separate per-modality pathways is tested by locating emotion-sensitive neurons — sparse decoder units selectively tied to emotion categories — in Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B, using speech emotion recognition and facial expression recognition as paired probes. Visual units turn out to be causal: silencing one degrades recognition of exactly its emotion, and amplifying it improves that emotion relative to others. Acoustic and visual units show emotion-matched overlap and similar layer distributions, and interventions transfer bidirectionally — units found in one modality produce emotion-specific effects when applied to the other — suggesting partly shared affective components that can be localized and steered without any training.
Which Source Wins? Task-Dependent Reliance in Vision-Language Models
When an image and its accompanying text support different answers, it is unclear how a vision-language model shifts its reliance as one source becomes harder to read. The experiments degrade either the image or the text across four legibility levels while keeping the other clean, building arithmetic conflicts by pairing a rendered GSM8K or SVAMP problem image with a different problem's text, plus ChartQA-Conflict, a new manually reviewed set of 229 chart-report conflicts with matched chart and table-image forms. Five of six open-weight models move away from degraded text on the arithmetic tasks, but all six show the opposite pattern on ChartQA-Conflict, abandoning the degraded visual source instead — a reversal that survives calibration for unimodal accuracy loss and replacing charts with plain tables, and that GPT-5.6-Luna and Gemini-3.5-Flash behaviorally reproduce.
Code as Representation: A Compilable Parsing Paradigm for Academic Documents
Academic PDFs interleave text with tables, formulas, charts, and pseudocode whose structure, data, and logic survive poorly in Markdown-style surrogates. CADP (Compilable Academic Document Parsing) reframes the task as representation rather than perception, reconstructing a full page as contextual LaTeX plus executable Python so the output can be recompiled and verified against the source page, and CADP-Bench supplies expert-verified full pages evaluated through a re-injection compilation protocol. Testing current multimodal large language models along with an exploratory multi-agent baseline shows that even frontier models fail to produce high-fidelity executable reconstructions.
Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models
Unified multimodal models (UMMs) are meant to let understanding and generation reinforce each other, but joint-training ablations cannot attribute gains to architecture rather than overlapping supervision. Binding a novel visual entity — a rendered 3D asset paired with a pseudo-word screened for absence from the frozen model's behavior — through exactly one direction and then measuring the untrained direction shows the channel is real both ways but asymmetric in kind: generation training installs a name the model can only match among candidates, while understanding training installs one it can also produce. What governs transfer is where the binding enters the shared computation; an alignment probe predicts export across 36 configurations at Spearman ρ = +0.68, the usable window appears only when the understanding pathway is a semantic vision encoder, and a mid-stack alignment objective acquires the concept for a 0.1% relative loss of general text-to-image ability against 41% for the standard generative route.
Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
Sustained conversation imposes demands that single-prompt evaluation misses: users clarify goals, revise requests, interrupt, switch topics, and add evidence while expecting context, memory, and grounding to persist across turns, tools, modalities, and cultures. Surveying text-only dialogue, audio LLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents, the authors organize the literature by datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. The central conclusion is that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session, with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment still unresolved, followed by a research agenda for systems that remember, revise, ground, speak, listen, act, and adapt.
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
Existing document benchmarks for multimodal models lean toward information extraction, often need external domain knowledge, and are dominated by English and Chinese. BEAR-Bench is a self-contained English-and-Russian benchmark of 1000 human-annotated questions over text-dense business and scientific documents, designed so that reasoning rather than retrieval determines the answer. Across 16 proprietary and open-weight models including Gemini 3.1 Pro and Qwen3.5-397B, clear headroom remains even for the strongest systems, and the collected outputs are reused to compare how reliably existing hallucination-detection methods catch those failures.
3 more specialized papers
- UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity Pengyu Wang, Baochen Xiong, Xiaoshan Yang et al.
- Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis Shanshan Lin, Yuesheng Wu, Chao Chen et al.
- SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis Shicheng Ma, Wenqian Cui, Irwin King
Reasoning 9
From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
Pipelines that have a language model translate a word problem into a formal specification and hand it to a solver break down when the translation is wrong but still runs, since existing repair loops only trigger on crashes. The proposed fix replaces the solver error message with a proof: when the generated program is unsatisfiable, a minimal unsatisfiable core over the model's own constraints is extracted and returned, pinpointing the exact set that cannot hold together without leaking the answer. On a new 77-problem benchmark with an exact oracle, translation to Answer Set Programming is faithful in six of seven domains, and the minimal core cuts a weaker model's rate of fabricating solutions to infeasible problems from 79% to 7%. A strong chain-of-thought baseline matches the symbolic route on accuracy, so the authors position the value as certificates and refusal to fabricate rather than raw correctness.
Solvable Sokoban Without a Solver via Diffusion
Deciding whether a Sokoban puzzle can be solved is PSPACE-complete, solutions can be exponentially long, there is no short certificate, and a single misplaced wall silently ruins a board — so generating solvable puzzles normally requires a solver in the loop. A transformer-based discrete diffusion model trained only on tile completion, with no solver, reward, or solvability label, instead produces solvable puzzles 77.4% of the time, with 94.5% of the remaining failures fixed by deleting one wall. The argued reason is order-freedom: an autoregressive model conditions each cell on a fixed prefix, whereas masked diffusion learns to fill any cell given any revealed subset, matching the non-local constraint structure that makes these puzzles hard. Training follows MD4 on DeepMind's Boxoban dataset, and the trained model plus generation instructions are public.
Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner
Claims about when a model "acquires" a capability usually rest on behavioral thresholds or on what a probe can decode from hidden states, and these need not mark the same moment. A 30-million-parameter recurrent-depth relational reasoner is studied in a closed oracle-defined world with dense behavioral trajectories, preregistered pre-arrival hidden-state probes, and untrained and negative controls. Three-hop competence took 70 logical epochs on a symbolic training surface but 13,055 on a verbal one — a 186.5-fold gap — after which four-hop verbal competence arrived in 8 more epochs, while during the long grind held-out four-hop accuracy never exceeded 3 of 40. Linear probes recovered future-answer identity before behavioral arrival on both surfaces (for example 0.056 versus 0.025 uniform chance and a 0.0248 untrained control, p = 0.013), but tracking that accessibility across training proved not cleanly evaluable because probe eligibility is itself defined by behavioral arrival, leading the authors to argue that only causal intervention can identify what training actually acquired.
The Problem Is the Problem: Towards Scalable Mathematical Discovery
In AI-assisted mathematics, expert time is consumed at two ends — choosing a worthwhile problem up front and reviewing generated proofs afterwards — and both are becoming bottlenecks as model reasoning gets cheaper. FAR (Find, Attempt, Recommend) replaces the single pre-chosen problem with a research direction: it mines a literature corpus for open conjectures, attempts them with frontier reasoning models, and triages results so only high-yield artifacts reach humans. A combinatorics pilot started from 5,245 papers, extracted 6,453 candidate conjectures, narrowed to 4,717 apparently open and well-posed ones, produced 598 potential resolutions, and surfaced 77 items for author review, including results on conjectures of Davies–Jenssen–Perkins–Roberts, Erdős–Straus, Ikenmeyer–Pak–Panova, and Lund–Saraf–Wolf.
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning
Post-training recipes that lift language models on competition math have seen little testing on signal processing, where problems mix symbolic derivation with domain conventions. Two pipelines are compared for adapting Qwen2.5-3B-Base to graduate-level problems from WirelessMATHBench-XL: reinforcement learning with verifiable rewards applied directly, and supervised fine-tuning on a distilled wireless chain-of-thought corpus followed by the same RL stage, each run with GRPO, GSPO, and GMPO. The strongest configuration reaches 39.12% accuracy against 12.37% for the untrained base model, roughly a threefold gain.
LLM-Only PDDL Domain Repair with Open-Weight Models
Automated planning depends on hand-written world models in the Planning Domain Definition Language (PDDL), and repairing faulty ones is typically framed as editing the model until it accepts given positive test plans and rejects negative ones. Recent open-weight language models were evaluated on this repair task with no symbolic machinery in the loop: the best model reaches an F1 of .87 against .49 for the symbolic baseline when given high reasoning effort. That headline hides a reliability problem — mean test pass rate in that setting is only .82 and collapses to .06 on the Thoughtful domain, and even the best configuration with test traces reaches .92, so these models cannot guarantee the constraints that automated repair requires.
Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning
Knowledge graph reasoning infers missing facts from graph structure, and while large language models do well at it through in-context learning, their parametric knowledge does not line up with the graph's structural context — a mismatch the authors name reasoning evidence perception drift, which hurts both accuracy and faithfulness. SIRLM (Structure-Internalized Rule Language Model) makes structural rule generation the central object, pairing an in-context learning block backed by a structural relation memory with a knowledge-graph tokenizer trained for structural invariance and a neuro-symbolic reasoner that propagates messages under rule constraints and feeds rule-execution results back as supervision. The approach drops into standard training pipelines such as supervised fine-tuning and GRPO, and outperformed 17 state-of-the-art knowledge graph reasoning methods across 36 datasets.
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing
What a reasoning model writes is only a partial record of the process producing it, so a two-level internal readout for mixture-of-experts models distills vocabulary-scale J-space into J64, a 64-axis semantic frame learned from the model's own reasoning states, which separates inference effort from problem-induced strain and adds 0.096 to 0.135 held-out AUC over a baseline reading the same rollout as token occupancy. R64 reconstructs J64 from native expert-routing statistics alone, reaching median per-axis correlation of 0.69 to 0.86 across three models and two families and preserving 95 to 100% of J64's predictive gain on gpt-oss-20b at low overhead. Both readouts drive test-time decisions: R64-weighted voting beats plain majority voting in seven of eight settings, and a rolling-window stop-and-resample policy with its operating point fixed on training questions improves accuracy by 1.1 to 5.9 points using J64 and 0.9 to 3.2 points using routing alone. Router edits targeting the mechanism J64 names induce the predicted behavior change, shifting a diagnosed stall from numerical guessing toward exact symbolic execution.
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning
Reasoning in language models is mostly measured in domains that hand the model the rules, such as mathematics and code; linguistic puzzles invert that, requiring the solver to first infer the system before reasoning inside it. The IOL-AI Challenge ran on unseen problems from the International Linguistics Olympiad (IOL) 2026 Individual Contest, attracting 731 submissions from 46 teams under a strict budget of one T4 GPU and 30 minutes, and was graded both automatically and — for the first time — by members of the official IOL Jury using the rubrics applied to human contestants. Among 15 additionally benchmarked unconstrained models, Claude Opus 4.8 earned a jury score equivalent to a gold medal, while the two resource-constrained systems submitted for jury grading scored in the bottom 5% of contestants; 14-billion-parameter entries beat models twice their size, with gains traced to decoding and output handling rather than capacity. Automatic metrics rank systems in exactly the jury's order but compress the scale, overscoring weak systems by about 13 points, and prior exposure to the puzzle languages did not measurably help frontier models.
Robotics 8
SkillComposer: Learning Reusable Skills for Natural-Language Robot Programming
Language models handle single-step robot commands but falter on multi-step tasks, high-level decomposition, and reusing solutions they have already found. SkillComposer pairs a generate-test loop, in which a model drafts and revises a robot program before execution, with an online library-learning algorithm that compresses recurring function sequences from successful programs into reusable macro skills available to later tasks. Ablations plus a 12-participant user study on simulated manipulation and robot caregiving tasks show that evaluator-guided generation together with the learned abstractions raises task success rates and usability while lowering the effort users spend specifying tasks in natural language.
NPU Offloading of a Frozen Visual Encoder for Robot Policy Training
Freezing a robot policy's visual encoder during training removes its backward pass, but the forward pass still runs at every step and keeps consuming GPU compute. The authors built an asynchronous pipeline that runs the frozen encoder of the AR-Actor specialist in INT8 on a Mobilint Aries2 neural processing unit (NPU) while the FP32 action expert trains on an RTX 5060 Ti, offloading between one and four Transformer encoder layers. Offloading all four cut energy per sample by 27.9% and peak GPU memory by about 20%, at the cost of 37.7% longer training time per sample. Across 4,500 simulator rollouts, success rate fell from 93.33% for GPU-only training to between 91.44% and 92.89%.
Teach and Grow: An Agent-Centered Architecture for General Robot Learning
End-to-end vision-language-action models break down when they encounter an object, sensor, embodiment, or contact outside their validated physical coverage, and fixing each failure requires new robot data, a policy update, and regression testing — what the authors call the retraining tax. Teach-and-Grow Learning instead has a multimodal agent turn a few successful demonstrations into reusable Skill Blocks, closed-loop behaviors for meaningful subgoals, which it grounds and composes in new scenes while a Skill Library and structured experience memory carry forward successes, failures, and repairs. The architecture reports state-of-the-art results on LIBERO without task-specific policy retraining, and the authors propose a scaling-law hypothesis in which future-task error and teaching demand approach irreducible floors as power laws in accumulated reusable experience.
5 more specialized papers
- Learning Varying Physical Therapist-Patient Interactions for Robot-mediated Upper Limb Task-Specific Training Jia Quan Loh (Human Robotics Laboratory, Department of Mechanical Engineering, The University of Melbourne) et al.
- ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai et al.
- tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots Markus D. Kobelrausch, Michael Miedler, Axel Jantsch
- Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision Amir Arsalan Nematollahi, Shayan Ahmadi, Mehdi Tale Masouleh et al.
- Dijkstra as an Oracle for Online Stochastic Shortest Path Navigation with Provable Guarantees Mansur M. Arief, Ali Akarma, Ahmad Alfan Alfian Irfan