Thursday, August 13, 2026
Highlights
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
Simulating societies of many large language model (LLM) agents is expensive, even though the questions asked are usually macroscopic, concerning phase behaviour and scaling with population size rather than any single agent's cognition. The method replaces each LLM agent with a low-parameter surrogate model fitted from a few hundred to a few thousand cheap queries, allowing the society to run at any scale on a laptop, and introduces a taxonomy crossing interaction order with memory that predicts in advance how surrogate error will trend with the number of agents. Validated on a faithful reimplementation of the EconAgent macroeconomy simulation and seven other named LLM simulations, with decisions cloned from genuine LLM elicitations for a few dollars, the predicted error trends hold cell by cell, and the two refuted predictions are themselves matched quantitatively by the theory with no free parameters.
Simulating societies of hundreds or thousands of LLM agents costs tens of dollars and hours per run, yet the questions researchers ask of them are almost always macroscopic — phase behavior, stylized facts, and how observables scale with population size N — making per-agent cognition mostly wasted expense. The proposed remedy is to replace each LLM agent with a surrogate of just 2–12 parameters, fitted by behavioral cloning from a few hundred to a few thousand cheap LLM queries, and then run the society at any N on a laptop; crucially, an "interaction order × memory" taxonomy predicts in advance whether this works, based on what each agent perceives: agents reacting to a global aggregate sit in a mean-field regime where surrogate error vanishes as N^(-1/2), while community-shared or graph-local feeds leave an O(1) error floor or worse, and a strongly curved response adds a Jensen-bias floor even under private signals.
Validated on a reimplementation of the EconAgent macroeconomy plus seven other named LLM simulations, with decisions cloned from genuine DeepSeek elicitations for a few dollars, the surrogate recovers EconAgent's Phillips correlation at −0.665 ± 0.12 against the published −0.619, while exposing that its celebrated Okun's law is an accounting identity that a behavior-free policy already satisfies. A 2×2 ablation further isolates the chain-of-thought reasoning step, not prompt wording, as what produces the emergent Phillips curve, and the taxonomy's two refuted predictions — both on a strongly saturating response — are themselves explained quantitatively by the measured curvature, whose Jensen bias of 0.018 matches the observed error floor with no free parameters.
The main caveats are that reproductions match mechanisms and scaling trends rather than published numbers from the targets' own models, the error decomposition is an organizing heuristic rather than a theorem, and results rest primarily on a single teacher model, with sibling models shifting the macroscopic values — a sensitivity the author frames as a finding in itself.
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
World modeling has no dominant recipe, since architectures, objectives, and state representations interact in complex ways, which makes it a natural testbed for AI coding agents doing open-ended research rather than implementing to a spec. The benchmark has frontier coding agents autonomously improve a provided world-model starter under a fixed compute budget across eight game environments, using ground-truth entity state in a shared tensor format so that dynamics modeling is isolated from perception and each run takes minutes. Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improved their starter in 63, and in 91% of sessions the winning edit was a research-style change, such as a new objective, representation, rollout procedure, or architecture, rather than a hyperparameter tweak.
Most benchmarks for AI coding agents hand them well-specified engineering tasks, leaving open the question of whether they can do actual research, where the direction of improvement is not given in advance — and world modeling, where no architecture or training recipe dominates, is a natural place to test that. The benchmark drops an agent into a closed loop with a starter world model (Dreamer-style RSSM, autoregressive Transformer, D3PM, or MaskGIT), a six-hour H100 budget, and eight arcade-style game environments whose ground-truth entity state is exposed in a unified tensor format, so the agent iterates on dynamics modeling directly without any perception stack and each training run completes in minutes.
Across 64 sessions, Claude Opus 4.6 and Codex-5.4 beat their starter on the held-out test split in 63, with a mean score lift of +0.196 on a 0-to-1 scale, and in 91% of sessions the winning edit is a genuine research-style change — a new objective, representation, rollout procedure, or architecture — rather than a hyperparameter tweak. The gains land almost entirely at long-horizon open-loop rollout (+0.215 at 20 steps versus +0.056 at one step), meaning the agents produce models that stay accurate under their own predictions rather than just fitting single transitions.
The two agents are statistically indistinguishable at this sample size (p=0.15), though Codex-5.4 reaches comparable scores with about 1.4x fewer tokens. The main caveats are that each agent-plus-harness is measured as a package, so model quality and scaffolding effects are conflated, and the state-centric setup means recipes discovered here may not transfer directly to pixel-based world modeling.
Agent Safety Should Be a Runtime Contract
The dominant paradigm treats AI safety as something instilled during training, which this position paper argues is structurally insufficient for autonomous agents that execute code, mutate files, send messages, and modify databases; safety should instead be a runtime contract enforced by the agent harness. The contract has a preventive face — sandboxes, permission gates, output filters, and trajectory monitors that block dangerous actions — and an evidential face that gates task submission on verifiable proof such as test runs, log captures, file diffs, and citation grounding. The position is supported by a survey of 52 documented agent and LLM safety incidents, a false-completion audit, a trajectory-schema audit of 12 public agent systems, and a title-level audit of 28,560 papers from NeurIPS, ICML, and ICLR 2023-2025 showing a pooled 8-12x publication imbalance favoring training-time over deployment-time safety; the authors formalize an Agent Trajectory Schema and Evidence Chain and outline a research agenda.
A position paper arguing that the field's dominant framing of AI safety as a training-time property — instilled via RLHF, DPO, or Constitutional AI — is structurally inadequate for autonomous agents that run shell commands, edit files, and touch databases, since a jailbroken or reward-hacking model leaves no backstop. The authors' alternative is a runtime contract enforced by the harness with two complementary faces: preventive mechanisms (sandboxes, permission gates, tool whitelists, trajectory monitors) that block dangerous actions, and evidential mechanisms that refuse to mark a task complete until the hash-chained trajectory contains verifiable artifacts such as test re-runs, file diffs, or citation lookups, formalized as an Agent Trajectory Schema and Evidence Chain with a compositional gating result.
Four empirical audits back the position: in a survey of 52 documented agent and LLM incidents, 40 are coded as fully preventable by a harness layer and only one as genuinely an alignment failure; a 32-case audit catalogs false completions like Replit's database deletion and the fabricated Mata v. Avianca citations; and only 2 of 12 audited agent systems document submission-time evidence gates, even though 11 of 12 already capture tool outputs.
A title-level analysis of all 28,560 papers accepted at NeurIPS, ICML, and ICLR from 2023 to 2025 finds training-time alignment work outnumbers deployment-time harness work by a pooled 8–12×, which the authors read as a research-allocation mismatch. The main caveats are that the contract only gates tasks with checkable acceptance criteria — open-ended creative work and mesa-optimization are explicitly out of scope — and the incident codings rest on counterfactual judgments while the proceedings audit is a keyword-based estimate rather than a full-text census.
Self-Evolving Embodied Agents via Skill-Harness Evolution
Adapting embodied agents to new environments typically requires fine-tuning or reinforcement learning, which demand extra data, rewards, and training runs, while training-free code-centric alternatives assume programmable robot APIs that fixed-interface settings lack. SHAPER keeps the foundation model's weights frozen and instead evolves the surrounding agent system — reusable skills and a context-code harness — through rollouts in the target environment, with the same frozen model serving as both planner and optimizer. On VLABench and ESI-Bench, spanning agents with different low-level action interfaces, this skill-and-harness optimization outperforms pure execution, supervised fine-tuning, and test-time-scaling baselines such as verifier-free selection and voting.
Adapting an embodied agent to a new environment typically means fine-tuning or reinforcement learning — which requires weights, data, and compute — or generating robot code against APIs that many fixed-interface settings simply don't expose. SHAPER instead keeps both the planner and executor frozen and evolves two external artifacts: a textual "skill" (procedural guidance like command formats, progress checks, and failure recovery) and a "context-code harness" (a Python function that selects and formats trajectory history for the planner), with the same frozen vision-language model doubling as the optimizer that rewrites both from summarized rollout feedback under beam search with sandboxed validation.
On VLABench, a frozen Qwen3.6-27B planner driving a π0 actor rises from 28.25% success as a seed agent to 34.5% after evolution, beating same-data supervised fine-tuning (24.0%) and test-time-scaling baselines like trajectory voting that actually fall below unscaled execution, with the largest gains (+6 to +10 points) on distribution-shifted splits. On ESI-Bench, an active spatial-perception benchmark, accuracy jumps from 32.5% to 49.8% micro, and the macro score of 42.9% nudges past a published GPT-5 single-view reference — though the authors flag this as an unpaired, external comparison on a 231-question subset.
The whole evolution run costs under $3 in API tokens and produces reusable artifacts, unlike per-episode sampling. Caveats worth noting: the skill and harness gains aren't additive, a few ESI-Bench categories regress under the evolved harness (some with as few as five questions), and everything is in simulation with real-robot validation left to future work.
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
Removing reflections from videos shot through glass has lagged behind the single-image case because paired video data, temporally coherent models, and benchmarks have all been missing. The S2R framework closes the loop: a physics-grounded synthesis pipeline augments structure-space representations with glass effects such as roughness-induced blur, thickness-induced ghosting, and reflectance variation, then renders realistic paired videos with a trained video diffusion renderer; the removal model adapts a pretrained video diffusion prior via reflection-aware latent adaptation to recover the clean transmission in a single denoising step. On the accompanying S2R-Bench and multiple public image benchmarks, the approach achieves state-of-the-art results while running faster than even non-diffusion baselines.
Reflections in videos shot through glass degrade quality and confuse downstream vision systems, but while single-image reflection removal is well studied, the video version has stalled for lack of paired training data, temporally coherent models, and benchmarks — applying image methods frame by frame produces flicker and unstable backgrounds. The authors, from Xiaomi's MiLM Plus lab, close the loop with three pieces: S2R-Synthesis generates aligned reflected/clean video pairs by compositing reflections in lineart structure space — with physics-grounded augmentations for roughness-induced blur, thickness-induced ghosting, and Fresnel reflectance variation — and rendering them photorealistically with a Wan2.1-based video diffusion model; S2R-Removal then adapts the same video diffusion prior via reflection-intensity supervision in latent space followed by one-step refinement with pixel, SSIM, and depth-consistency losses, recovering the clean video in a single denoising step; and S2R-Bench provides the first dedicated evaluation suite, pairing 60 full-reference videos with 50 in-the-wild clips scored by humans.
The one-step design pays off doubly: 28.84 dB PSNR on the benchmark versus 27.04 for the best frame-wise diffusion baseline, near-perfect human scores for transmission preservation (0.98), and 87 ms per frame — about 1.67× faster than the quickest non-diffusion competitor and roughly 80× faster than the diffusion one. A data ablation shows the synthesis pipeline is independently useful, lifting an existing image-domain model by 1.78 dB over planar-only compositing.
The caveats are that the full-reference benchmark is built from static image pairs animated with virtual camera motion rather than genuinely dynamic captures, real-world evaluation rests on 15 human raters, and the authors acknowledge weakness on multi-layer reflections and geometry coupled to camera motion.
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
Agentic systems for business ideation have so far been text-only, despite real business contexts being inherently multimodal. MBA-Bench provides 30K samples across six domains with distinct visual cues, using automatic image captioning and GPT-4o to generate reference ideas through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis, with evaluation by a multimodal LLM judge across six business criteria. Two agents are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with creativity and feasibility rewards — MBA-b for when evaluation criteria are hidden and MBA-k for when they are disclosed — and they outperform caption-only baselines by 63.9% and 77.1% and multimodal baselines by 25.6% and 35.8%, respectively.
Business ideation agents have so far worked in a text-in, text-out paradigm grounded in patent documents, even though real-world opportunities often hinge on visual details — a crowded storefront, a subtle surface defect, a cluttered app layout — that captions demonstrably fail to convey. The authors build MBA-Bench, 30K image–caption–question–idea samples spanning six domains chosen for varying verbalizability (everyday scenes, mobile UI screenshots, crowds, anomaly images, textures, and circuit boards), with reference ideas generated by GPT-4o through retrieval of market evidence via DuckDuckGo.
On top of the benchmark they train two 7B agents from Qwen2.5-VL via LoRA fine-tuning followed by GRPO: MBA-b for when evaluation criteria are hidden, optimizing only creativity and feasibility rewards, and MBA-k, which additionally optimizes the six disclosed judging dimensions; feasibility is scored not by a judge but against a web-sourced grounding library using FAISS retrieval and FActScore-style fact checking. The agents beat caption-only baselines by 63.9% (MBA-b) and 77.1% (MBA-k) and multimodal baselines by 25.6% and 35.8%, with the 7B MBA-k matching or exceeding GPT-5 and Gemini-class models on innovativeness, competitive advantage, and market size under an MLLM judge.
The main caveat is that everything — reference ideas, training rewards, and final scores — flows through large-model judges with no direct human evaluation, so the headline numbers measure agreement with an InternVL2.5-78B rubric rather than validated commercial merit, and results come from a single training run.
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
Large language model (LLM) agents that call external tools are vulnerable to indirect prompt injections hidden in environment state, but existing security studies rely on hand-built environments and predefined injection locations. ToolHazard is a scalable framework that synthesizes executable, stateful adversarial environments using an Environment Simulator, an Attacker Agent, and a User Simulator, which together discover viable injection points, generate environment-specific payloads, and construct long-horizon tasks. The resulting ToolHazard-Bench reveals substantial agent vulnerabilities and shows that injection timing and placement affect attack success, while alignment data generated by the framework improves security on both ToolHazard-Bench and AgentDojo without degrading benign task performance.
Tool-using LLM agents are exposed to indirect prompt injection—malicious instructions planted in emails, database records, or tool outputs the agent reads mid-task—but studying this at scale has been bottlenecked by hand-built environments with predefined injection points. The authors instead have LLMs synthesize the entire adversarial testbed: an environment simulator generates executable, stateful tool environments as verified code (at roughly $0.59 per environment), an attacker agent automatically discovers writable-and-readable injection points and plants payloads, and a user simulator produces long-horizon benign tasks, with success checked programmatically against the final environment state rather than by an LLM judge.
The resulting benchmark spans 28 environments, 512 tools, and 87 tasks averaging 15.6 steps, and frontier models fare poorly—four of six attack strategies exceed 40% attack success rate against GPT-5, and DeepSeek-V3.2 pairs the best benign task completion with the worst vulnerability, suggesting stronger instruction-following can mean easier hijacking. Two structural findings stand out: injections encountered earlier in a trajectory and placed near the end of a tool observation succeed more often, and free-form text outputs are markedly more exploitable than JSON or YAML.
Fine-tuning Qwen3-8B with SFT plus GRPO on 1,040 synthesized adversarial samples roughly halves attack success on both the in-distribution benchmark (36.1% to 18.1%) and the independent AgentDojo (29.2% to 18.3%) while improving benign completion. The catch is that everything is synthetic—environments may miss production-system failure modes, the six injection strategies are fixed rather than discovered, and alignment was only validated on small open models.
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Turning a research idea into a complete paper requires literature retrieval, experiment design and execution, evidence-based revision of claims, publication-ready figures, and consistency over a long generation process. The system implements this pipeline as thirteen composable skills inside an existing coding assistant, separating model-based judgment from deterministic checkable operations, specifying required evidence before results are observed, and bounding a failure mode in which repeated experiments keep rejecting the original research objective. Across eight controlled research topics it achieves 99.5% citation validity and 96.4% figure editability, raises fabrication detection from 14% for a single-pass draft to 92% with the full integrity and review stack, and averages 11.9 million tokens, $8.10, and 3.2 hours per manuscript.
Autonomous research agents like AI Scientist can run the full ideation-to-manuscript cycle, but they ship as standalone platforms with their own orchestration servers, separate from the coding environments where research actually happens. The authors instead implement the entire pipeline as thirteen composable skills inside Claude Code, splitting each task between model judgment (organizing arguments, assessing evidence) and deterministic scripts (citation validation, LaTeX compilation, result-integrity gates), with experiment designs and result-table structures committed before any results are observed so that claims are revised against measured evidence rather than fitted to it.
Two mechanisms stand out: a bound on what the authors call Self-Refutation Loops, capping experiment-critique-revision cycles at seven and converting unsupported research directions into failure reports rather than forcing them into success narratives, and a figure pipeline that plots quantitative results programmatically from data while reconstructing image-model-generated method diagrams into editable vector PDFs via iterative HTML rebuilding. Across eight pre-registered topics the system achieves 99.5% citation validity and 96.4% figure editability, and an ablation with 36 injected unsupported claims shows fabrication detection climbing from 14% for a single-pass draft to 92% with the full gate-plus-review stack, at an average cost of 11.9M tokens, $8.10, and 3.2 hours per manuscript versus $0.66 and 16 minutes for the single-pass baseline.
The comparisons with prior systems rest on audits of their publicly released artifacts rather than reruns under matched conditions, the evaluation covers only eight topics, and the crucial step of judging whether evidence actually supports a claim remains a model decision rather than a machine-checked one.
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
As AI development becomes faster and more automated, mechanistic understanding of models remains largely manual, widening the gap between what models can do and our ability to understand and control them. The proposed agentic system autonomously discovers mechanisms underlying model capabilities, supported by an interpretability-focused knowledge graph of about 13,000 papers, a 43-million-paper multidisciplinary database, and a curated library of 32 methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems it generates more valuable mechanism hypotheses and executes experiments more reliably; its reported discoveries include a safety risk in which unsafe traits transfer across modalities through apparently safe training data, a mechanistic theory of how models represent knowledge and infer beliefs, and interventions that improve performance and steer scientific foundation models toward generating DNA sequences with specified properties.
Mechanistic interpretability is still hand-built work: model capabilities are advancing under increasingly automated research pipelines, while the effort to explain those capabilities stays manual, and the handful of existing automated interpretability tools mostly generate and check descriptions of individual neurons or features rather than proposing general mechanism theories. Mechanist is an agentic system that runs the full loop — hypothesis generation, experiment execution, verification, and iteration under a central orchestrator, with each agent in an isolated context passing explicit file artifacts downstream — grounded in an interpretability knowledge graph of roughly 13,000 papers and research blogs, the 43-million-paper SciAtlas corpus across 26 disciplines for cross-domain analogies, and a library of 32 executable mechanistic methods covering probing, activation patching, circuit discovery, and sparse autoencoders. On a benchmark of 16 recent interpretability papers where each system receives only the target claim and is denied access to the paper or its repository, Mechanist scored highest on reproduction reliability under all three judges (three human experts, Claude Opus 5, and GPT-5.6-sol) and produced hypotheses rated more novel, impactful, and testable than those from Claude Code and AI Scientist. Its three case studies are the more interesting evidence: it extended subliminal learning to the multimodal, semantically-opposing case, where a student fine-tuned only on text responses that a GPT-4o filter judged safe reached a 48.6% unsafe-response rate on multimodal lab-safety questions against 20.3% untuned and 18.3% for a regular-teacher control, and a student trained on banana-free apple images from a banana-preferring teacher generated bananas 25.6% of the time versus 2.5% and 2.1% for the baselines. It then localized separable belief circuitry in Pythia — zeroing the attributed-belief head L4.H1 dropped attributed-belief accuracy from 0.86 to 0.34 while personal-belief accuracy held at 0.71 and Pile perplexity moved only from 7.96 to 8.05 — traced those heads' emergence across pretraining checkpoints, and converted the finding into a probe-gated inference-time intervention worth +15.3%, +8.8%, and +3.5% on Pythia-410M/1B/2.8B against +1.6%, +3.1%, and +0.1% for oracle-style prompt hints; a parallel application steered a sparse-autoencoder feature in Evo2-7B to raise mean predicted α-helical content across 900 generated DNA sequences from 43.8% to 56.6%, holding up under pLDDT filtering while random-feature steering did nothing.
The caveats are worth weighing against the framing. The belief mechanism work is confined to Pythia and OLMo at 410M–2.8B parameters because those are the families that publish intermediate pretraining checkpoints, so the developmental account is untested at frontier scale; the biological results are ESMFold predictions rather than experimental structures, and the steering coefficient sweep shows control degrading into invalid open reading frames past α=8; reliability scoring leans partly on LLM judges; and the authors themselves recommend running the system as a human-AI co-scientist rather than autonomously, which concedes that the verification agent is not yet a sufficient check on its own conclusions.
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Instead of distilling a large model's capabilities into a smaller one through parameter updates, this work tests whether transfer can happen entirely at inference time: a stronger builder model designs a harness, meaning scaffolding code and prompts wrapped around a weaker target model, without any retraining. On four Theory-of-Mind benchmarks, the builder iteratively refines its harness using 5% of the data as a validation set, and the finished harness nearly doubles the target model's average score from 0.49 to 0.91. Analysis attributes the gains to offloading unstable reasoning into deterministic code, benchmark-specific routing, and strict answer-format enforcement rather than to more extensive reasoning or sampling by the target model, with the weakest targets benefiting most and harness quality improving monotonically with the builder's reasoning effort.
Instead of closing the gap between strong and weak language models through distillation or fine-tuning, the authors ask whether a strong "builder" model can transfer capability purely at inference time by writing a harness — prompt templates, benchmark routing, deterministic solvers, and format enforcement — around a frozen weaker target. The builder sees only a 5% validation slice and must ship an executable scaffold that generalizes to a 3,900-item hidden test set spanning four Theory-of-Mind benchmarks.
Across 57 runs, every builder configuration beat the baseline, with a mean macro-accuracy uplift of +0.275 and a best run (GPT-5.5 building for GPT-5.4-mini) that lifted the target from 0.49 to 0.91 — surpassing the unscaffolded, much larger GPT-5.4 at 0.62. The gains track builder quality and reasoning effort rather than validation-set probing (best validation score predicts test accuracy at r=0.96, while the number of refinement iterations is nearly uncorrelated at r=0.17), and the strongest scaffolds work by compiling task structure into code — polarity logic, belief-state extraction, deterministic offloading — rather than by eliciting longer reasoning.
Uplift follows a "headroom law": the already-strong Gemini-3.5-flash target gained only +0.11 and actually regressed on benchmarks near ceiling, a caution that scaffolding can interfere as well as assist. Automated scaffolds also still trail a human-designed harness (0.939) on the harder benchmarks, with residual errors concentrated in deep recursive belief tracking and Bayesian goal inference, and part of the BigToM gain reflects exploiting a compilable shortcut rather than genuine reasoning support.
Applications 97
Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet
Reinforcement-learning post-training now dominates language-model development, yet its GPU power behavior is uncharacterized, and datacenters manage power with workload-blind static caps and reactive throttling. The authors instrument GRPO training at 7B, 14B, and 72B scales with half-second power telemetry (over 380,000 samples) and train a PPO meta-controller that adapts the workload's own generation parameters to measured power, cutting power-limit violations by 89.8% on the 7B trace while increasing token output by 18.1% and energy efficiency by 26.2%. At 72B the original actuator lost authority under model sharding, but a controller rebuilt on generation concurrency delivered 35.7% more output than a static safe baseline with 87.2% fewer violations than uncontrolled operation across three live replications. A composed 16-GPU fleet analysis shows zero violations at 30-second measurement windows with peak demand at 50-56% of nameplate, suggesting roughly twofold power oversubscription is feasible for this workload mix.
Why AI Detection Fails for Academic Integrity
Institutions use commercial AI detectors as evidence in academic integrity cases, yet detectors cannot distinguish light AI editing from fully LLM-drafted text. In a controlled study of published English abstracts across four domains (2013-2015 versus 2023-2025), lightly refined abstracts — a proxy for guideline-compliant AI assistance — were flagged 64-80% of the time by Pangram and GPTZero, while 9-15% of unmodified recent originals were also flagged, with non-STEM fields flagged far more often than STEM. After processing with a commercial humanizer tool, fewer than 4% of AI-labeled rewrites remained detectable, meaning honest AI-assisted editing carries higher sanction risk than deliberate evasion; the authors conclude detector scores should not serve as standalone misconduct evidence.
Methodologies for Improving the Quality of AI Tutoring in K-12 Education
AI tutors built on large language models (LLMs) are hard to improve reliably because the underlying models are opaque, making rigorous evaluation and live experimentation essential for measuring the impact of every change. Drawing on experience running Khanmigo, Khan Academy's K-12 AI tutor launched in 2023, the authors describe the metrics they use to measure tutoring quality and student engagement along with the experiments they have run. They highlight the changes that measurably moved those metrics, including model choices, prompting, personalization, and agents.
Long-Horizon Forecasting of Complete Financial Statements with Forma
In a discounted-cash-flow valuation most of a firm's value lies beyond the next year, yet no prior work jointly forecasts complete financial statements past that window. ProForma-20Q is released as a reproducible benchmark for predicting 78 statement line items one to twenty quarters ahead for anonymized firms, given only past statements and an industry code, scored by change-space R-squared. Forma, a transformer that reads statements as sets of (account, quarter, value) tuples and maximizes a masked-tuple Gaussian likelihood, outperforms classical machine learning, chained gradient boosting, a zero-shot time-series foundation model, and frontier large language models, with its margin widening at longer horizons and its Gaussian predictive intervals never under-covering. Its forecasts nearly satisfy accounting identities, exact coherence can be imposed at no statistically significant accuracy cost, and the tuple interface supports scenario analysis without retraining, with pinning a future revenue path sharpening the rest of the statement.
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs
Estimating many non-commuting observables of a quantum state under a finite measurement budget forces a trade-off between statistical efficiency and circuit resources like depth and entangling-gate count, with existing strategies clustering at the extremes of shallow product measurements or deep fully commuting ones. FlowMeas recasts resource-constrained measurement design as generative learning, using a generative flow network to directly sample ensembles of shallow Clifford measurement circuits under a prescribed shot budget and hardware constraints. At zero entangling depth it matches or beats leading product-measurement methods on nearly all molecular benchmarks, one or two entangling layers reduce energy-estimation error by up to 27% over the strongest state-independent product baseline, learned policies transfer across related Hamiltonians, and the framework scales to a compactly encoded 54-qubit interacting fermionic model.
Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling
Drug response prediction (DRP) models for anticancer therapy are limited by small datasets, narrow coverage of cancer and chemical spaces, and inconsistent benchmarking. The authors greatly expand the benchmark from the Innovative Methodologies and New Data for Predictive Oncology Model Evaluation (IMPROVE) project by integrating large-scale pharmacogenomic data, primarily from PharmacoDB, adding millions of drug response measurements, broader multi-omics coverage, and more than 50,000 new compounds. Models trained on the expanded dataset match the original benchmark on cancer-blind splits but consistently improve on drug-blind and disjoint splits, indicating better generalization to previously unseen compounds.
Probing and steering biology across Boltz-1s trunk-diffusion boundary
Structure predictors in the AlphaFold3 family pair a representational trunk that processes sequence and context with a diffusion module that generates atomic coordinates, and it is unclear how biological information changes as it crosses that boundary. Linear probes, sparse autoencoders (SAEs), and causal interventions applied to per-residue activations of Boltz-1 show that geometry (secondary structure, disorder) and sequence chemistry (amino-acid identity, signal peptides, disulfide-bond annotations) are both linearly decodable from the trunk, but inside the diffusion module secondary structure transfers nearly unchanged while sequence chemistry is strongly attenuated. Steering the final trunk representation that conditions diffusion changes predicted structure dose-dependently for helix and coil directions, yet a highly predictive beta-strand direction (F1 = 0.82) produced no measurable increase in strand content, so linear decodability does not imply causal influence at that site. The authors also note that probe scores against sparse SwissProt annotations are lower bounds, because correct predictions on unannotated residues are charged as false positives, and they release the trunk and diffusion SAEs, activations, and analysis code.
RECAST: A Machine-Learning Framework for Correction and Super-Resolution of Coarse-Grid PDE Solvers
Coarse-grid solvers make time-dependent partial differential equation (PDE) simulation much cheaper, but under-resolution degrades both the solution trajectory and its spatial detail. RECAST (Recurrent Error Correction And Super-resolution of coarse-grid Trajectories) inserts a learned correction inside the numerical time-stepping loop and reconstructs the fine-grid state from the corrected coarse history. Across six one-dimensional PDE systems with grids coarsened by factors of 8 to 16 and 1000-step closed-loop rollouts, it cuts time-averaged relative error by roughly 50-92% versus uncorrected coarse solvers, generalizes to unseen PDE parameter values, and beats a contemporary coarse-correction architecture on 5000-step rollouts.
Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models
Telling dengue-infected mosquitoes apart from controls by their movement in video is difficult because the insects are tiny and the backgrounds are complex, which trips up conventional feature extraction. A three-step pipeline first detects mosquitoes and removes the background with a YOLO 11M model, extracts visual features with a Vision Transformer (ViT), and classifies the videos with a convolutional gated recurrent unit (ConvGRU). In a comparison against recurrent neural network, long short-term memory, and plain GRU variants, ConvGRU performed best at 88.88% accuracy and an 82.81% F1 score, suggesting that combining convolutional feature extraction with sequence modeling captures both fine spatial detail and long-term movement patterns.
CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications
Automatically tagging daylong audio recorded in infants' homes is hard because labeled data is scarce, signal-to-noise ratios are low, and acoustic conditions shift from family to family. The system combines a Whisper encoder fine-tuned with low-rank adaptation (LoRA) with a lightweight target-speaker-aware Transformer that performs long-context, framewise prediction across annotation tiers, adds a sequence-level smoothing loss for temporal coherence, and uses a factorized speaker-token design—a shared tier token plus a learned family-specific offset—to reduce family bias. Together these choices enable efficient and robust infant-centered tagging of naturalistic daylong home recordings.
Making AI-Generated Feedback Matter: From Provision to Student Enactment
Generative AI can deliver individualized feedback at scale, but students' uptake of that feedback remains low. A large-scale quasi-experimental cohort study spanning 13,037 students and 51,296 student-authored resources compared three workflows: AI feedback comments delivered without support, optional student-initiated AI dialogue, and an enacted workflow that prompts students to select suggestions, evaluate their relevance, and discuss them in targeted dialogue. The enacted condition achieved an estimated 26.2% probability of feedback uptake versus 14.1% for direct provision and 0.1% for optional dialogue, along with higher self-assessment confidence and submitted-work quality, indicating that workflow design — not just feedback quality — determines whether AI feedback improves learning.
Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians
Computing ground states of quantum Hamiltonians beyond the reach of classical simulation is a central target for quantum computing; this work instead amortizes the problem across arbitrary quadratic qubit Hamiltonians with a roughly 0.5-billion-parameter foundation model trained using techniques from large language models and deep reinforcement learning. Spin-1/2 ground-state learning is reformulated as variational optimization over functions on the SU(2)^N manifold, with the Hamiltonian acting through Lie derivatives, and the variational principle is proven to preserve the ground-state upper bound. The model is pre-trained on hundreds of thousands of Hamiltonians of varying topology, size, and interaction type using a replica-exchange Langevin sampler and an extended Kronecker-Factored Approximate Curvature (KFAC) optimizer on systems up to 64 qubits, then fine-tuned on held-out systems up to 1024 qubits and evaluated on systems up to 8100 qubits.
Task- and dataset-specific information in protein language models
Protein language models (PLMs) are typically used by taking embeddings from their final layer, even though little is known about what intermediate layers encode. The authors probe every layer of 13 PLMs across 15 downstream tasks drawn from 11 datasets, training probe models per layer and characterizing the latent spaces. They find that the last layer rarely yields the best embeddings: residue-level tasks resembling the pre-training objective improve steadily with depth, while for whole-protein tasks the dataset matters more than the task itself, with deep mutational scan data favoring shallow layers and diverse natural proteins favoring deeper ones; performance also drops significantly on artificial proteins.
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
Recent claims that general-purpose language models have caught up with specialized clinical AI rest on a narrow set of comparisons and on benchmarks built largely in high-income settings. VITA, a retrieval-augmented generation (RAG) system whose curated corpus covers disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols, was scored against frontier models on 4,023 English HealthBench questions using a GPT-4.1 judge. It placed first with 51.9% of available rubric points, ahead of GPT-5.4 at 46.1%, and on a 500-question rerun graded by an open-weight judge sharing no lineage with any system tested it was statistically indistinguishable from GPT-5.5 on mean per-question score while leading on points-weighted score. Its advantages came from accuracy and completeness; communication scores were lower.
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
Protein structure prediction models still fail on some targets, and the external biological oracles that can flag and correct those failures are expensive, making the number of oracle calls the binding constraint. The authors benchmark guidance methods that spend that budget in different ways — FK-steering, direct preference optimization (DPO), and best K-of-N sampling — alongside Optimisation Over Outputs (O3), which runs off-the-shelf optimisers within a generative model's latent subspace and which they extend to structure prediction. On calmodulin (1CLL) and E. coli aspartate transcarbamoylase (9EEH), no method dominates across budgets and oracles: O3 is strongest at low oracle budgets while FK-steering and DPO improve as the budget grows, which the authors turn into budget-dependent recommendations for practitioners.
ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening
Drug combinations lower the risk of resistance to any single agent, but the search space makes exhaustive combinatorial screens prohibitively expensive and often technically infeasible, and existing predictors need molecular profiling of each sample plus per-cohort training. ScreenShot is a hierarchical transformer, pretrained on 40 drug screening datasets covering 3,700 drugs and 6,000 biological samples, whose architecture mirrors the nested structure of screening data; given a few-shot context of observations from a new patient sample it predicts combination responses by in-context learning, working directly on functional measurements with no fine-tuning and no molecular profiling. It outperforms all baselines on four held-out datasets for both prediction accuracy and identification of selectively effective treatments, and its internal representations drive a weighted k-means++ active learning strategy that matches uniform screening's hit detection using a third of the experimental budget.
How Organizations Use AI: Evidence from ChatGPT
Linking ChatGPT Enterprise account records to usage logs, worker roles, task classifications, and public-company financial data through March 2026, the study analyzes enterprise AI adoption at scale, covering over 1,500 organizations and 17 million messages at the six-month adoption horizon. It documents four facts: usage has grown through both new firm adoption and rising intensity among existing adopters; U.S. public-company adoption concentrates in larger, more valuable, R&D- and SG&A-intensive firms; use spans job functions and seniority levels with early-career workers the most intensive users; and messages cover a broad range of knowledge work including writing, technical tasks, communication, and information synthesis. The authors conclude that firms differ widely in the speed and breadth of adoption and are still learning how to integrate AI into workflows.
80 more specialized papers
- Evaluating LLM Generated Detection Rules in Cybersecurity Anna Bertiger, Bobby Filar, Aryan Luthra et al.
- Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration Patrik P. S\"uli, Gy\"orgy Eigner, Roland Holl\'os
- LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs Timur Zakarin, Sergei Voitov, Sergei Shumilin et al.
- Transit Destination Inference from Tap-In-Only Bus Smart-Card Data: A Hierarchical Bayesian Approach Gefei Zhao, Jiahe Ling, Yuelong Su
- Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction Jiaquan Zhang, Shuxu Chen, Haifan Meng et al.
- BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model Jia-Rui Lin, Junxi Guo, Keyin Chen et al.
- Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach Chaofan Zhai, Yicheng Song, Ravi Bapna et al.
- FarSky: Task-Aware Latent-Space Coupling for Generative Intra-Hour Solar Forecasting Yann Fabel, Bijan Nouri, Milon Miah et al.
- Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures Bongseok Kim, Suman Chakraborty, Gary Huang et al.
- Temperature-Driven Sequential Modeling for the Prediction of Annual Power Conversion Efficiency Profiles of Organic Photovoltaic Materials: Douala Case Study Steve Cabrel Teguia Kouam, Rockefeller Rockefeller, Raoult Dabou Teukam et al.
- CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data Fenosoa Randrianjatovo, Maya Saleh, Simon Girard et al.
- Hardware-Aware Deployment of Joint SAR Compression and Despeckling on FPGA C\'edric L\'eonard, Francescopaolo Sica, Martin Schulz
- Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification Rofiqul Islam, Lilatul Ferdouse
- Federated Learning for Distributed CNC Tool Wear Prediction Afsana Khan, Morris Stallmann, Marcin Pietrasik et al.
- Physics-Informed Implicit Neural Representations for Improved Myocardial Perfusion MRI Quantification Christos Tsepas, Chang Yan, Maximilian Fuetterer et al.
- Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models Guobin Zhao, Xiao-Yan Li
- Benchmarking Cyberattack Detection in Electric Vehicle Charging Infrastructure with Benign User Updates Hannan Chen, Roshni Anna Jacob, Jie Zhang
- Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin Detection Using Texture Analysis and Deep Learning Pouya Afshin, Tianling Niu, Tongtong Lu et al.
- Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces Mengyu Chen, Feiyu Lu, Chun-Fu Chen et al.
- Market-Information-Aware Gated-LoRA of Foundation Models for Transferable Day-Ahead Electricity Price Forecasting Hang Fan, Wei Wei, Shengwei Mei
- Towards an approach to multivariate outlier detection for District Heating System data Rajko Turudija, Du\v{s}an Stojiljkovi\'c, Milan Zdravkovi\'c et al.
- From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate Pardis Taghavi, Santosh Bhavani
- Diffusion-Based Data-Driven Assortment Optimization Junyi Liao, Xiaohui Jiang, Zhengwei Tong et al.
- Stigma and Support in Online Sexual Violence Narratives on Reddit Shirlene Rose Bandela, Karan Bindal, Vaibhav Garg et al.
- Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates Qiyao Zhou, Xujia Zhu, Pierre Joli et al.
- DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition Akriti Dhasmana, Aarohi Srivastava, David Chiang
- XGBoost "is all you need": the case of forecasting transmitted heat energy in District Heating Systems Milan Zdravkovi\'c
- Gaussian Meta-Space Augmentation for Stacking Ensembles in Multimodal IPMN Risk Stratification Max A. Nelson, Eminenur Sen Tasci, Zhixiang Wang et al.
- Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware Sadib Hassan Rumman, Md. Shariful Islam, Md. Rayhanur Rahman
- FLARE++: Low-rank attention with dynamic attention routing Vedant Puri, Yongjie Jessica Zhang, Levent Burak Kara
- Hierarchical Federated Transfer Learning in Digital Twin-Based Vehicular Networks Qasim Zia, Saide Zhu, Haoxin Wang et al.
- Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough Sang Su Lee, Vineeth Loganathan, Shishir Dash et al.
- RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation Yueyuan Li, Zexi Chen, Weijie Xi et al.
- FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting Rentao Gu, Yihang Ding, Junjie Li et al.
- Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones Oshan A. B. Yalegama, Wageesha N. Manamperi
- Easper: An Accessible ASR Pipeline for Language Documentation Aso Mahmudi, Ting Dang, Ekaterina Vylomova et al.
- Transferable Above-Ground Biomass (AGB) Estimation Model from Multi-Sensor Data with Sparse Field Calibration Pann Thinzar Seint, Bryan Atwood, Subas Chhatkuli
- Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Huaxuan Wang, Huimin Wang, Ruiyu Zhang et al.
- FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation Yu Zhang (AMap Alibaba Group, Beijing, China) et al.
- AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection Touseef Hasan, Mounika Ghanta, Souvika Sarkar et al.
- LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification Michael Schlee, Fabian Lukassen, Christoph Weisser
- MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques Jiabao Zhuang, Changhao Jiang, Hanchen Wang et al.
- Automated binary classification of hazelnut X-ray images: A deep-learning benchmark for quality assessment Giancarlo Sportelli, Nicola Belcari, Roberta Pace et al.
- A comparison of CNN architectures for Alzheimer's disease detection in single-view MRI scans Hiram Zuniga, Ulises Orozco-Rosas, Kenia Picos
- Instruction Alignment for Binary Code Representation Learning Huaijin Wang, Shuai Wang
- Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines Vaishnav Raju
- TradingMoE: Routing the Right Experts in Evolving Markets Chang Zhou, Xingtong Yu, Minbin Huang et al.
- High-Order Liquid Evidence Encoding for Gradual GNSS Spoofing Detection in Autonomous Driving Muhammad Ayub Sabir, Junbiao Pang, Fatima Ashraf
- JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time Series Yian Wei, Yuanyuan Yao, Lu Chen et al.
- Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly Detection, Attack Identification, and Hardware Monitoring Martin Sachenbacher, Martin Leucker, Alexander Weiss et al.
- Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs Alireza A. Safaei, Laura M. Vowels, Matthew J. Vowels et al.
- Air Quality Station Simulation via LSTM and Attention-Based Modelling Alexander Kostadinov, Petar O. Hristov, Dessislava Petrova-Antonova
- User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magn\'usson et al.
- When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation Jinhyung Bae, Dain Kil, Seongmin Oh et al.
- Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra Waleed Waseer, Muhammad Shahid Jabbar, Muhammad Sohail Ibrahim et al.
- Forward and Inverse Virtual Metrology for Phototransistor Gain: A Hierarchical, Uncertainty-Aware Approach for Small Production Datasets Mahshid Amirabgir, Lorenza Ferrario, Paolo Conci et al.
- DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation Anik Pramanik, Murat Kantarcioglu, Vincent Oria et al.
- LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training Xiaojun Wu, Cehao Yang, Honghao Liu et al.
- Distillation of Foundation Models for Time-dependent PDEs Daniel Musekamp, Boshra Ariguib, Andrei Manolache et al.
- TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement Karim Aly, Alexei Sharpanskykh, Jacco Hoekstra
- HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs Kangning Zhang, Haotian Fang, Xukun Luo et al.
- A Remote Approach to Cashew Orchard Detection: Leveraging Active Learning with Satellite Imagery in Guinea-Bissau Miguel, Sofia, Maria et al.
- Beyond Local Power: Functional Connectivity Analysis for Subject-Independent Learning Style Recognition Wiga Maulana Baihaqi, Indriana Hidayah, Sri Kusrohmaniah et al.
- Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh Muhammad Masud Tarek, Md. Alamgir Hossain, Md. Samiul Islam et al.
- Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches Muntasir Hasan Kanchan, Md. Alamgir Hossain, Md. Samiul Islam et al.
- Poly-Dialectal Neural Machine Translation System for Bangla Regional Dialects Rakib Ullah, Ruhul Islam Rahul, Tanbir Ahmed
- From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices Tuhinangshu Gangopadhyay, Rasmus Adler, Peter Liggesmeyer et al.
- Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP Nikolette Pedersen, Regitze Sydendal, Veronika Cheplygina et al.
- RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation Rong Chao, Sung-Feng Huang, Moreno La Quatra et al.
- Attractor Image-Based Deep Learning of Arterial Pulse Waves for Age Classification Sara Vardanega, Patrick Segers, Philip Aston et al.
- FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees Zhiqiang Que, Chang Sun, Haiyang Wang et al.
- Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification Kazi Nabiul Alam, Pooneh Bagheri Zadeh, Akbar Sheikh-Akbari
- Few-Shot Ordinal Learning for Day-Wise Freshness Estimation with Hyperspectral Fish Images Kazi Nabiul Alam, Pooneh Bagheri Zadeh, Akbar Sheikh-Akbari
- VICBench: A Multi-Language Benchmark for Code Vulnerability Detection Jin Lu, Xuening Han, Yang Zhong et al.
- Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting Junyi Ye, Gargi Vijay Borde
- Calibration Bets on the Past: Post-Training Quantization for Financial Time-Series Forecasting Junyi Ye, Ivy Gateri Wanjiku
- A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement Bryan Torres, Daniel Riofr\'io, Jos\'e Vega-S\'anchez et al.
- Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling Pedro Sousa (Department of Computer Science, University of Cambridge), Will Tebbutt (Department of Engineering et al.
- Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals Alireza Kargarzadeh, Nariman Khaledian, Navid Parvini et al.
- Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models Saman Marandi, Yu-Shu Hu, Mohammad Modarres
Large Language Models 48
What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model
Methods that feed a language model its own output, such as self-consistency, iterated refinement, and agentic loops, raise the question of what such probes actually measure. The authors study this in a deliberately sharp setting: a ring of token cells resampled in place by the model's own windowed conditionals, with a coupling scheme that makes damage spreading between perturbed copies exactly measurable. They find the readings split into two look-alike kinds: quantities fixed by the construction itself, such as the damage light cone and the radius scaling of a token-space Lyapunov exponent that is invariant across 19 models, and quantities that genuinely track the model, such as where that exponent crosses zero during training. They report having confused the two themselves for four months, including a precisely measured phase transition that belonged to the probe rather than any model, and propose a test: hold the construction fixed while varying the model, or vice versa, and see which readings move.
Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts
In top-k Mixture-of-Experts (MoE) models, routing is discontinuous, so numerical noise from 4-bit key-value cache quantization can push tokens across decision boundaries and flip which experts fire. Rather than proposing a fix, the work builds a causal measurement apparatus that quantifies the route-mediated fraction (RMF) of quantization damage, finding on OLMoE-1B-7B that roughly a third of the damage flows through routing changes, a result that carries across three architectures via pre-registered probes. A deployable router-margin statistic detects that a flip occurred (AUC 0.772) but performs at chance when distinguishing harmful flips from helpful ones, and no tested inference-observable statistic predicts a flip's effect on loss above chance, establishing an empirical barrier to selectively repairing routing damage at inference time.
Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction
When context windows fill up, LLM systems compact prior context to continue a task, and user-issued session constraints, such as an instruction not to delete any emails until the user confirms, are silently dropped in the process. The evaluation suite measures this loss across multi-turn chat, agentic trajectory, and long-horizon research scenarios, finding that current compactors retain only 17% of injected constraints on average and that most perform worse than running the same task without compaction. Retention varies sharply with compactor, prompt, context length, constraint phrasing, and injection location, indicating the loss is systematic rather than tied to one setting. A constraint-aware extractor running alongside the compactor as a plug-and-play module achieves over 90% retention across all three scenarios without modifying the compactor or the model.
Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier
In collaborative multi-model settings, peer opinions compete with a model's own parametric knowledge, and measurements across 23 open-weight models, 19 conditions, and over a million graded responses show a unanimous wrong majority reverses 22.8% of correct MMLU answers, 54.8% on GPQA, and 71.0% on SimpleQA. Existing mitigations target Resistance, the rate at which a model keeps a correct answer under pressure, which the authors pair with Receptivity, the rate at which a model adopts a correct peer answer after initially erring. Scoring six mitigation methods on both axes shows each gains Resistance only by losing Receptivity, with method means falling on a single trade-off frontier at an R-squared between 0.80 and 0.90. Reasoning is the sole exception: on MMLU subjects whose answers a model can derive for itself, it raises Resistance by 7.2 points and Receptivity by 9.6 points at once.
Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport
Adapting a language model to each individual author by supervised fine-tuning (SFT) requires separate weight access, optimization, storage, and retraining per person, which makes personalization prohibitively expensive. Weightless Fine-Tuning (WFT) skips training entirely: it computes supervised residuals on an author's training sequence and transports them to the current prompt through a cross-prefix transport operator estimated from dropout-induced cross-covariance, so gradient-based parameter updates are replaced by corrections in logit space. On three LaMP personalization benchmarks WFT achieves the best average performance, matches or exceeds SFT on individual tasks, beats other lightweight baselines on average, and in a budget-controlled comparison approaches SFT using under 7% of the effective computation. The logit shifts it induces have cosine similarity 0.875 with those from SFT over 95% of the next-token probability mass, suggesting it reproduces the distributional effect of supervised adaptation without touching weights.
Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter
Tokenizer vocabulary size for large language models (LLMs) is usually fixed at training time by convention, but the cost-optimal choice actually depends on the serving regime. The authors formalize lifecycle cost as training cost plus inference volume times serving cost, and run controlled experiments on A10G and A100 GPUs spanning memory-bound and compute-bound regimes. They find the inference-optimal vocabulary shifts 16-fold with serving batch size — from about 32k at batch size 1 to 524k at batch 64 and above, driven by amortization of the unembedding matrix read — while quality measured in bits per byte varies under 2% across the optimal range, making vocabulary size effectively a pure systems parameter. Their guidance: roughly 32k vocabularies for on-device serving and 131k-262k for datacenter serving.
Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models
The lack of diversity in language model outputs is widely blamed on alignment, but where in the training pipeline the collapse actually begins has been unknown. Through controlled supervised fine-tuning (SFT) experiments, the authors find that training data can reveal and amplify semantic convergence but not introduce it, suggesting SFT acts as a catalyst rather than a cause. They further show that instruct-like output collapse can be induced in base models through prompting alone, without any alignment, indicating that homogeneity likely arises from the pretraining objective itself and is difficult to fix with post-alignment interventions.
Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task
Congruency effects in conflict tasks such as Stroop and flanker have been studied in psychology for nearly a century without a settled mechanistic account. The authors build a verbal-only conflict task for language models, where a prompt stem invites a default same-color completion and an explicit rule either agrees with it (congruent) or contradicts it (incongruent), and test Gemma-2-2B plus six Pythia models from 410M to 12B parameters; all showed strong default same-color tendencies and six of the seven showed large congruency effects. Causal attribution, attention analysis, and attention ablations expose two pathways: short-range attention to a superficial color cue that is preferentially engaged on congruent trials, and long-range attention to the rule prefix engaged on incongruent ones. Fine-tuning that strengthened the default mapping improved congruent and degraded incongruent performance, while enlarging the rule set selectively hurt incongruent trials, supporting an account in which the effect arises from competition between an in-weight default mapping and an in-context rule.
Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation
Prompt wording and structure are known to affect large language model (LLM) output, but psychologically motivated framings had not been studied for coding tasks. Building on Yukl and Falbe's taxonomy of influence tactics, the authors turn eight strategies including rational persuasion, ingratiation, and exchange into reproducible prompt templates and evaluate them on five leading open-weight models using LiveCodeBench and SWE-bench Verified, scoring functional correctness, code quality, maintainability, and security. Framings that emphasized urgency were associated with lower correctness and weaker security. The study is presented as the first large-scale empirical look at influence-induced prompt framing in software engineering, with practical guidance for designing transparent human-AI interaction in code generation.
Unifying Physical Backpropagation
In enterprise retrieval-augmented generation (RAG) deployments, large language models satisfy about 80% of individual constraints yet only 26.8% of responses meet all requirements simultaneously—a 57-point orchestration gap that existing benchmarks, which assume clean retrieval and simple queries, fail to capture. The benchmark comprises 983 expert-validated samples across six domains that systematically combine complex multi-dimensional instructions with three production failure modes: retrieval noise, knowledge gaps, and factual conflicts. Evaluating 13 state-of-the-art models reveals a severe collapse in holistic instruction adherence, with knowledge gaps and factual conflicts remaining hard even under reasoning-enhanced inference, pointing to the need for explicit context-aware protocols and calibrated judgment in production RAG systems.
CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement
When user queries are ambiguous or incomplete, large language models (LLMs) that answer directly tend to produce overgeneralized or erroneous responses, and existing methods for learning when to ask clarifying questions depend on costly human annotation. CLAIM removes that dependency by quantifying query uncertainty through the entropy of answer disagreements across multiple models, combining this signal with semantic clustering and reasoning-based judgments to automatically synthesize training data, then training a unified clarification decision model with supervised fine-tuning followed by group-relative policy optimization. Experiments show the model learns stable, generalizable strategies for deciding both when clarification is needed and which aspect of a query to clarify, without manually labeled data.
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
Unstructured knowledge editing injects a free-form passage of new facts into a large language model, but existing editors leave the model able to recall the passage without answering atomic questions about its facts or composing them into multi-hop reasoning, a missing property the authors call composability. They recast editing as proactive self-distillation from a privileged in-context state of the same model, and observe that pure on-policy distillation struggles because the pre-edited model's rollouts rarely cover the novel knowledge. Their method, HPSE, builds hybrid rollouts that insert missing facts into the student's own trajectory exactly where coverage fails while staying on-policy elsewhere, with a theoretical analysis of the advantage and empirical gains across four backbones and two editors.
The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
Benchmark scores rest on a single phrasing of each problem, and rephrasing a problem while preserving its meaning and answer routinely flips model answers in both directions, turning failures into successes and vice versa. BenchDrift generates meaning-preserving variations along linguistic, referential, pragmatic, and structural axes and measures how often correctness flips under each. Across eight models on GSM8K, MMLU, and MATH-Hard, phrasing sensitivity does not fade with capability but changes sign: weak models gain more from rephrasing than they lose while strong models lose far more than they gain, and models largely agree on which rephrasings are costly, indicating the fragility belongs to the rephrasing rather than the model.
Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models
Diffusion large language models promise faster inference through parallel decoding, but existing schedulers commit each position only when it independently meets a confidence criterion, ignoring how early commitments can help later ones. The authors identify a ripple effect in which proactively committing a mid-entropy pivot position sharply reduces uncertainty across remaining masked positions, and propose a training-free decoder that selects such pivots and picks their token assignments via lookahead evaluation. Across three diffusion language models and four reasoning and code-generation benchmarks, the method achieves 4-10x wall-clock speedup over standard decoding while preserving quality, improves accuracy over a prior lookahead baseline by up to 5.49%, and reaches up to 18x speedup with key-value caching.
Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization
Which training data helps models generalize to tasks never specified during training? Building on epiplexity, a recently proposed measure of the structural information a compute-bounded learner can extract from data, the authors turn it into an online training signal: for data selection they fit scaling laws to per-domain loss curves to predict epiplexity gain and adaptively reweight domain sampling, and for synthetic data generation they use REINFORCE policy gradients to reward a generator for increasing learner epiplexity. In both settings, higher epiplexity predicts better downstream zero-shot and fine-tuning performance, supporting the hypothesis that structurally rich data yields representations that transfer across domains.
Hybrid Gated Attention
Gated attention mitigates attention sinks and improves the representational capacity of attention; this work pushes its effectiveness-efficiency frontier with a framework combining three gating strategies that draw on information from multiple stages of attention and build element-wise and head-wise gates capturing both intra-head and cross-head interactions. Low-rank matrix decomposition and a learnable attention sink are added to improve training efficiency and stability. Across backbones and standard benchmarks the approach improves training loss and downstream performance over gated attention and achieves the best results at different computation costs.
Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release
Many interpretability studies report that language models encode task-relevant latent structure they fail to use, but whether located structure can actually be converted into behavior is rarely tested end to end. The authors run a fully preregistered stress test of a detect-localize-release pipeline on a 25.7M-parameter transformer trained on causal-evidence discrimination, with every threshold and decision rule hashed before data collection. Localization succeeds: interventions at mid-layer observation-evidence channels restore suppressed behavior with release advantages of 0.563 and 0.854. However, the gating detector inverts out of distribution, firing on 6.9-7.3% of generations that do not need it and on none of the 2,400 that do, and unconditional linear release plateaus far above the preregistered sufficiency threshold, showing the whole family of linear release directions at this site is bounded away from working.
Small-Scale Experiments: Are We There Yet?
Researchers have found scaling laws unreliable at small model sizes and concluded that sizable models cannot be avoided in experiments. The authors argue the confounding factor is hyperparameters: small models are highly sensitive to them, and scaling laws only emerge on the fully tuned frontier, which requires far more extensive search than most studies run. They show well-tuned hyperparameters matter more than any other ingredient of the scaling-law recipe and that the hyperparameter loss surface becomes lower-dimensional as scale increases, making tuning easier for large models. Synthesizing these insights into a methodology for small-scale research, they recover a known large-scale result, that pre-normalization outperforms post-normalization as transformers grow, from small-model experiments alone.
Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
Multiple-choice benchmark scores conflate a model's knowledge with its sensitivity to the order of answer options, making them unreliable measures of knowledge. The study tests whether preventing a model from seeing option labels while committing to an answer removes positional influence, comparing a generation-then-matching strategy with scoring each option in isolation, which is positionally unbiased by construction. Neither reliably improves accuracy: a full decomposition shows the bottleneck is withholding the options rather than the matching step, eliminating positional influence entirely still does not reliably help, and cyclic permutation of options often improves accuracy where debiasing does not.
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
Training large language models on limited hardware is increasingly a scheduling problem spanning GPU compute, host memory, PCIe transfer, and storage bandwidth, and existing offloading systems still leave communication exposed on the critical path. LazyTrain adds an optimization layer over a layer-streaming executor that formulates checkpoint selection, activation placement, recomputation, and CPU-GPU-NVMe communication overlap as a mixed-integer scheduling problem, and couples 8-bit optimizer states with fast gradient clipping in a single hybrid operator. Across H800 experiments from Qwen2.5-3B to Qwen3.6-27B it improves sustained TFLOPS by about 1.24 times over matched baselines, reaching 219.95 TFLOPS at batch size 72 on the 27B model, and on an RTX 3090 it raises the maximum feasible batch size at every model scale.
RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
Current benchmarks for large language model (LLM) generation of Triton GPU kernels restrict tasks to PyTorch-to-Triton translation, measure isolated kernel speed rather than end-to-end performance, and rely on hand-written evaluation scripts whose flaws models can exploit for inflated scores. This benchmark instead derives generation tasks from real pull requests that modified Triton kernels in popular open-source AI frameworks, pairing each natural-language requirement with a concrete engineering context and a reproducible evaluation environment, and validates generated kernels by integrating them back into their original frameworks and running end-to-end tests. Evaluations of leading LLMs show they still struggle with these production-like kernel generation tasks.
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges
Rubric-based evaluation with large language model judges usually treats the rubric as flat prompt context, leaving how criteria combine implicit even when the rubric's natural-language rules spell it out. Graph-Structured Rubrics (GSR) compile a rubric into a response-independent typed evaluation graph before any responses are seen: criterion nodes elicit judgments, transformation, reduction, and gating operators compose them through named ports, and a readout maps the graph's sink to a score or preference, with compilation rejecting malformed or type-incompatible graphs. Using GPT-OSS-120B as the judge, GSR improves exact score agreement by 0.62 to 6.75 percentage points over Prometheus-style scoring on four pointwise datasets and posts the highest end-to-end pairwise accuracy on two preference benchmarks.
QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving
Retrieval-augmented generation (RAG) serving repeatedly prefills the same text chunks across queries, and while position-independent caching (PIC) of key-value states avoids this, its efficiency is limited by the volume of text tokens; rendering chunks as images compresses them into fewer visual tokens but degrades quality due to contextual mismatch between independently compiled caches and loss of fine-grained text. QV-PIC compiles visual caches offline under the model's native chat-template prefix and, at query time, keeps global context at low resolution while spending a high-resolution budget on the chunks most relevant to the query. Across six tasks it improves average F1 by 21.6 points over vanilla rendered-image PIC, beats optimized text PIC by 2.58 F1 while cutting time-to-first-token by 17.2%, and reduces time-to-first-token by 83.8% relative to full prefill.
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
Some hidden units in transformers carry activations orders of magnitude larger than the rest, and this study tracks those massive activations in hybrid linear attention language models that interleave linear attention with full attention layers. Two architecture-aligned shapes recur: spikes immediately before each full attention layer, and plateaus in which a spike's magnitude persists across the intervening linear attention layers, with denser full attention linking successive spikes into the continuous pattern familiar from full-attention models. The organization holds across five linear attention architectures, six hybridization configurations, five data domains, and open models from 1.2B to 397B parameters, and controlled pretraining up to 1.3B shows that full-attention output gating sharply reduces magnitudes without disturbing the layerwise structure, while removing the gated-delta-network gates amplifies them only modestly. The authors attribute the two shapes to when the activation gets cancelled: spikes are written, used as a sink, and cancelled locally, whereas plateaus reflect delayed cancellation.
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
Leaderboards implicitly assume model rankings hold regardless of how many tokens a model is allowed to generate. Varying that maximum output budget across seven levels from 64 to 4,096 tokens for four models on three reasoning benchmarks, totalling 56,476 inferences, breaks the assumption: rankings reverse across budgets on every benchmark, and 3-19% of items become less accurate as the budget grows even after controlling for truncation, with only 6-14% overlap in which items misbehave across models. An oracle choosing the best model per item gains up to 27.8 percentage points, most of it at constrained budgets, but a practical budget-aware router recovers just 14.1% of that gap and its budget features help within a domain while hurting transfer across domains. The authors conclude that evaluation protocols should report results conditioned on generation budget.
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
Training with ever longer contexts is usually assumed to be free upside, since the model simply sees richer evidence. The authors argue the opposite can occur: when relevant information is already abundant in the training context, the model has less incentive to encode it in its parameters and instead learns to rely on context, a trade-off they call the Information Abundance Paradox. In pretraining on long documents, language modelling, natural language understanding, and closed-book multiple-choice accuracy improve with context length only up to an intermediate optimum and then decline consistently; in supervised fine-tuning, more task-relevant training context helps when supporting context is present at test time but reduces robustness when it is absent or misleading. Mechanistically, informative context shifts gradient pressure from feed-forward networks toward attention modules, and causal interventions confirm this shift increases context reliance at inference.
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Instead of distilling a large model's capabilities into a smaller one through parameter updates, this work tests whether transfer can happen entirely at inference time: a stronger builder model designs a harness, meaning scaffolding code and prompts wrapped around a weaker target model, without any retraining. On four Theory-of-Mind benchmarks, the builder iteratively refines its harness using 5% of the data as a validation set, and the finished harness nearly doubles the target model's average score from 0.49 to 0.91. Analysis attributes the gains to offloading unstable reasoning into deterministic code, benchmark-specific routing, and strict answer-format enforcement rather than to more extensive reasoning or sampling by the target model, with the weakest targets benefiting most and harness quality improving monotonically with the builder's reasoning effort.
21 more specialized papers
- MaSRead: Content-Addressed Reading of Replicated Latent Stores Carlos Baquero, Lu\'is Brito, Jo\~ao Resende
- From Monolithic to Modular: Segment-level Automatic Prompt Optimization Nikita Kulin, Viktor Zhuravlev, Artur Khairullin et al.
- LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs Yirui Liu, Ruoling Qi, Longwen Wang et al.
- CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference Yifan Wu, Yufeng Zhang, Kenli Li
- TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation Jiahui Zhang, Ziwei Zhang, Yipeng Wang et al.
- Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability Jeonghwan Choi, Taewon Yun, Minjeong Ban et al.
- Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression Angelo Nardone, Paolo Ferragina
- From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation Alireza S. Ziabari, Kat Ellis, Colleen Chan et al.
- Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models Yoshihiko Kayama
- APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference Alish Kanani, Layan Badawi, Umit Y. Ogras
- REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Yang Sun, Lichao Ma, Houyuan Qin et al.
- Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages Nirmal Thomas
- TELLME: Test-Enhanced Learning for Language Model Enrichment Minjun Kim, Inho Won, Hyeonseok Lim et al.
- Orientation, not magnitude: the causal structure of task-vector interference in merged language models Chencheng Zhu
- Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework Avinash Agarwal, Vridhi Jain
- LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu et al.
- Asymptotic Risk Calibration for Selective Question Answering Shufan Lin, Sijin Dong
- SoftWater: Class-Aware Rate Allocation for Softmax Quantization Joao V. Cavalcanti, Ashia C. Wilson
- Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations Lior Baruch, Moshe Butman, Kfir Bar et al.
- SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges Yuchao Wu, Junqin Li, XingCheng Liang et al.
- NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation Jiarui Ma, Jianghan Wang, Yuheng Ma et al.
Agents 42
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
Simulating societies of many large language model (LLM) agents is expensive, even though the questions asked are usually macroscopic, concerning phase behaviour and scaling with population size rather than any single agent's cognition. The method replaces each LLM agent with a low-parameter surrogate model fitted from a few hundred to a few thousand cheap queries, allowing the society to run at any scale on a laptop, and introduces a taxonomy crossing interaction order with memory that predicts in advance how surrogate error will trend with the number of agents. Validated on a faithful reimplementation of the EconAgent macroeconomy simulation and seven other named LLM simulations, with decisions cloned from genuine LLM elicitations for a few dollars, the predicted error trends hold cell by cell, and the two refuted predictions are themselves matched quantitatively by the theory with no free parameters.
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
World modeling has no dominant recipe, since architectures, objectives, and state representations interact in complex ways, which makes it a natural testbed for AI coding agents doing open-ended research rather than implementing to a spec. The benchmark has frontier coding agents autonomously improve a provided world-model starter under a fixed compute budget across eight game environments, using ground-truth entity state in a shared tensor format so that dynamics modeling is isolated from perception and each run takes minutes. Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improved their starter in 63, and in 91% of sessions the winning edit was a research-style change, such as a new objective, representation, rollout procedure, or architecture, rather than a hyperparameter tweak.
InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk
AI agents are a promising route to automating increasingly complex computing infrastructure management, but how well they handle real-world complexity has been unclear. The benchmark suite evaluates agents on realistic infrastructure tasks spanning the full system stack and operational lifecycle, with fine-grained risk assessment through per-check scoring. Across 15 agent-model configurations, mean effective scores range from roughly 40% to 88%, no agent achieves a full score, and repeating each task three times shows even top configurations pass only a fraction of attempts. A general failure pattern emerges: agents routinely satisfy short-term objectives while leaving behind non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state.
RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle
Deploying large language model agents in industrial recommender operations exposes a trilemma among general autonomy, industrial determinism, and end-to-end efficiency, where maximizing any two comes at the expense of the third. The platform, deployed for 78 days across three heterogeneous Tencent recommender business lines, grants autonomy only at bounded decision points inside pre-committed pipelines: an event-driven runtime consumes zero CPU during the 94% of wall-clock time spent waiting on Spark or GPU jobs, a 29-file skill ecosystem compiles per-skill pitfall tables into a 400-entry store that confines agent decisions, and a human-in-the-loop card protocol keeps operators at the diagnostic-versus-execution boundary. Over the deployment window the platform recorded 1,624 command-line tool dispatches with a 78.6% aggregate success rate, and onboarding-time compression was observed on two of the three business lines as a case-study observation.
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
Enterprise buyers read agent leaderboards as rankings of capability, and this analysis argues they largely rank specialization instead. A four-facet Generalizability Theory variance decomposition of three open agent-trace benchmarks (TheAgentCompany, tau-squared-bench, and AppWorld), fit with three estimators that agree to three decimal places, attributes under 3% of total variance to the agent itself and 7-23% to the agent-by-task interaction. Four further results sharpen the picture: aggregate reliability collapses on the hardest task quartile (from 0.752 to 0.000 on one check type), designs that look most reliable in the training cells replicate worst (correlation of -0.90), population-level diagnostics transfer across benchmarks while per-family rankings invert, and failure-mode profiles from the MAST taxonomy generalize at the cell level but not per trace. The authors package this into Deployment Decision Reliability, a one-page reporting discipline that turns the variance-component table into five procurement decisions, releasing code, data loaders, and fit artifacts openly.
Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost
Augmenting language-model agents with learned skills has become common practice, but existing work measures performance gains rather than cost, leaving unclear which skill-learning strategy is actually cheaper. The argument advanced here is that treating skills as programs saves the most, because deterministically executing a known action sequence avoids the trial and error and long-horizon degeneration that would otherwise be needed to reach the same goal. SpeedRunner tests this as a coding agent that analyzes its own past trajectories at inference time and refactors them into reusable programs, on the hypothesis that trajectories alone carry enough signal without replay or separate validation. Across three embodied environments it is reported to sit at the frontier of both task learning and cost reduction while remaining robust to distribution shift and environmental randomness.
Self-evolving network verifiers
Symbolic network verifiers can reason about correctness across vast spaces of routing inputs and failure scenarios, but only for the protocols and features an expert has hand-encoded, and keeping that model faithful never ends because vendor implementations deviate from the RFCs and behaviour shifts between software releases. The proposal is to grow the model automatically against the one unambiguous specification of what a network does, the router software itself: in a counterexample-guided loop, a coding agent proposes extensions to the verifier's symbolic encoding while a trusted oracle such as emulated routers supplies ground-truth routing state, and each disagreement drives the next refinement. A prototype taught a 3,000-line SMT-based verifier three features it did not previously support, OSPF areas, BGP route reflection, and L3VPN over EVPN, converging autonomously on models matching the oracle and even capturing vendor-specific behaviour. The authors note this shifts the hard problem from writing verification systems to systematically testing them, and outline a research agenda for trusting automatically evolved verifiers.
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Frontier models can solve hard tasks once the problem, tools, and success criteria are specified, but consequential real-world challenges rarely arrive in executable or verifiable form. Apodex Discovery addresses that gap through a "heavy-duty solver" combining a foundation model, harness, tools, and control policies to pursue extended, stateful, verifiable investigations, and it has three parts: a problem-scouting process that surveyed 561 industries across 16 sectors, assembled 423 high-value real problems, and released 20; a shared environment-task-episode abstraction supplying data, tools, constraints, feedback, trajectory recording, and verification of both intermediate artifacts and final submissions; and HDS6, which scores Tools, Repair, Alternatives, Coherence, Evidence, and Scope separately from final-task success. Reported results include exceeding the published state of the art by 7% on adeno-associated virus (AAV) capsid design across viability, tropism, structure prediction, and generative design, and raising mean normalized prediction scores for two GPT-5-series models by 2.5 and 7.6 points on a biomedical drug repurposing and reformulation environment relative to the same closed-book backbone. Controlled ablations indicate the fixed episode interface lets performance differences be attributed to specific solver components.
Self-Evolving Embodied Agents via Skill-Harness Evolution
Adapting embodied agents to new environments typically requires fine-tuning or reinforcement learning, which demand extra data, rewards, and training runs, while training-free code-centric alternatives assume programmable robot APIs that fixed-interface settings lack. SHAPER keeps the foundation model's weights frozen and instead evolves the surrounding agent system — reusable skills and a context-code harness — through rollouts in the target environment, with the same frozen model serving as both planner and optimizer. On VLABench and ESI-Bench, spanning agents with different low-level action interfaces, this skill-and-harness optimization outperforms pure execution, supervised fine-tuning, and test-time-scaling baselines such as verifier-free selection and voting.
Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology
Multi-agent systems have been proposed for medical diagnosis with large language models (LLMs), but it has been unclear when they are actually needed and where they outperform a single model. The authors build a multi-round pipeline that structures interaction among specialist agents as a deliberative process modeled on clinical differential diagnosis, and evaluate it against single-agent baselines, one-agent pipeline ablations, and best-of-n scaling. The multi-agent approach achieves a recall advantage that monolithic inference does not reproduce, with the largest gains on the hardest cases, where repeated rounds of specialist conversation help recover ground-truth diagnoses.
Benchmarking LLM Judges for Mobile Agent Evaluation
Mobile agent benchmarks increasingly use large language model (LLM) judges to decide whether tasks were completed, but the reliability of those judges has gone largely unexamined. The authors assemble 931 human-annotated agent trajectories spanning six mobile benchmarks, four agent models, and 68 apps, and use them to evaluate six judging methods across multiple LLM backends. A simple baseline judge given sampled screenshots matches or beats purpose-built judging pipelines, with the LLM backbone mattering more than pipeline design; judge quality metrics predict both agent-ranking fidelity and usefulness as reward signals for on-policy reinforcement learning; and different backends show qualitatively opposite failure profiles, one conservative and one permissive.
The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark
Much of the software that matters most for cybersecurity, including malware and firmware, exists only as binaries, and evaluating AI agents on reverse engineering (RE) them is hard: benchmark programs must be absent from training data to prevent recognition shortcuts, while still matching real software's scale and anti-analysis protections. The authors spent over 5,000 expert hours building 19 private, realistic programs averaging 16.9 thousand lines of code, combined with 44 in-house anti-analysis primitives to produce 262 binary instances and 1,572 deterministically graded tasks. Across five frontier models, the strongest scores 61.4% per instance and fully solves only 31.5% of instances, agents prove insensitive to compiler optimization and static linking in ways human engineers are not, and ablations confirm that both contamination control and realistic scale materially change results.
A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization
Hit-to-lead work in drug discovery requires repeatedly designing analogs of a hit compound under competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic-accessibility constraints. SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration) is an open-source framework in which a large language model interprets natural-language objectives and routes tasks to specialized tools for reaction-template analog enumeration, physicochemical and ADMET property prediction, structure-based affinity scoring, and Bayesian optimization, recording provenance for each numerical output. The resulting workflow mirrors the analysis and prioritization stages of the design-make-test-analyze cycle, and across single- and multi-objective studies it enriches candidate sets for the stated computational objectives while evaluating only part of the enumerated search space. Tools and characterization backends can be swapped by editing a config file rather than changing the orchestration logic.
Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents
Uncertainty quantification (UQ) methods for language models are usually validated on single generated answers, but large language model (LLM) agents produce interactive trajectories in which errors from tool calls and intermediate decisions propagate to the final outcome. The study tests whether three families of single-turn methods transfer to this setting—white-box scorers based on action-token probabilities, black-box consistency scorers over resampled trajectories, and reflexive self-assessment—across five LLMs and four multi-turn tool-use datasets from BFCL-v4 and tau^2-bench. Transfer proves useful but uneven: token-probability scores are highly sensitive to how per-turn values are aggregated, reflexive scores are the strongest low-cost baseline, and black-box self-consistency (especially trajectory-equivalence and action-set variants) is often the strongest family overall, leading the authors to argue that single-turn UQ methods need revalidation at the trajectory level.
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
Mobile graphical user interface (GUI) agents tend to break on applications absent from their training data, and adapting to a new app must happen within a limited interaction budget and without target demonstrations. CoAdapt-GUI is a test-time adaptation (TTA) framework that jointly updates two things from the agent's own rollouts and rewards: a structured workflow context that retains transferable procedures, failure modes, and verification rules while excluding app-bound details, and a policy, adapted via group-relative optimization of a LoRA adapter on a frozen vision-language model. It reaches 45.0% on AndroidWorld-Generalization versus 37.5% for a policy-only adaptation baseline, and raises AndroidWorld Plus performance from 38.6% to 52.9%.
Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem
Memory systems are widely deployed in large-model agents, yet there is no unified formal account of what a memory is or when it is optimal. The proposed framework treats memory as a basis and knowledge as its span: an agent stores events, a generation operator expands them into entailed knowledge, a query is answerable exactly when some item in the span covers it, and the optimal memory is the capacity-constrained maximizer of expected coverage, tracing a utility-capacity frontier on which systems can be compared. The account extends to noisy memories, where coverage diverges from precision, and formalizes continual memory construction as a sequential Markov decision process with memory as state, writing as action, and query-time utility as delayed reward; the framework is instantiated concretely on Homer's Odyssey and used to position existing memory systems.
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication
Multi-agent systems built from different large language model families can outperform homogeneous ones, but existing communication either goes through text, discarding internal representations, or requires identical architectures. The authors identify an entity grounding problem in cross-architecture latent transfer: continuous cross-attention bridges lose entity identity through rare-token compression collapse, scoring only around 30% F1 on their own. XBRIDGE addresses this with Lexical Anchor Mapping, which maps the sender's context tokens into the receiver's vocabulary as discrete entity anchors, plus a Latent Enrichment Bridge that lets the receiver query the sender's hidden states; across Llama, Qwen, and Mistral pairs on seven benchmarks it beats text-based communication on every task with 11x lower latency, using only 264M trainable parameters.
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents
Existing finance benchmarks mostly test data extraction, a narrow task current models have largely saturated, and reference-based or generic judge scoring handles open-ended analyst queries poorly. This benchmark comprises 220 expert-crafted queries with 11,543 source-attributed rubrics spanning six use cases across the full investor workflow, evaluated under a common harness restricted to public data. Results show the tool harness, not the model alone, strongly shapes quality: Samaya's in-house system leads at 56.0% versus 49.2% for the strongest frontier model (Claude Fable 5) at roughly 2.2x lower cost, the best open-weight model (Kimi K3) reaches 46.4% at 4.5x lower cost, and Screening & Discovery plus Sector, Industry & Macro remain hardest, with top systems at only 33% and 39%.
When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use
Large language models calling application programming interfaces (APIs) in multilingual settings often pick the correct tool but generate argument values in an inconsistent language, a failure termed Argument Language Mismatch that is operationally invalid yet invisible to standard API-calling metrics. Revisiting post-training strategies, the authors find supervised fine-tuning is a strong baseline that substantially improves argument language consistency and end-to-end call accuracy, matching or sometimes exceeding more complex reinforcement learning approaches under consistent model selection. Reinforcement learning with argument-aware rewards, such as Group Relative Policy Optimization (GRPO), adds only incremental gains, mostly in generalization and multi-objective trade-offs.
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
When a coding agent obeys a rule, existing benchmarks cannot tell whether it actually complied or would have behaved that way anyway. The benchmark scores operational rules one at a time from execution evidence across 60 multi-turn coding tasks, placing rules on the five configurable surfaces a deployed agent reads, and introduces Against-Prior Accuracy, which counts only rules opposing a model's unprompted defaults as observed by re-running tasks with the rule withheld. Across 12 frontier models, overall accuracy spans 72.1-85.9% but against-prior accuracy is 3.6 to 7.4 points lower for every model, so aggregate scores overstate compliance; a conflict pilot also finds precedence does not follow prompt depth, with system prompts, project files, and user instructions winning over tool and skill descriptions.
Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction
Coding agents self-correct well because compilers and tests turn failures into typed recovery signals, but broad language-agent tasks expose only coarse failure, and generic recovery playbooks widen the agent's context exactly when a narrower repair interface is needed. DARC, a diagnosis-guided recovery harness, uses development-set failures to profile a task family's failure modes, prunes mismatched interventions from a shared recovery library, and freezes a verifier-selected success-cost policy before deployment, so the agent first determines what kind of failure is repairable and then decides how much recovery evidence to spend. On ALFWorld, AppWorld, and XBRL Finance, the same protocol improves average task performance over base agents and broad playbooks while reducing environment steps or retrieval budget.
The Sleeping Agent: What Gist-Based Context Compression Loses and Why
Summarizing older conversation history into compact gists is common in long-horizon language model agents, but its effect on different kinds of memory retrieval is poorly understood. Using Salience-Weighted Consolidation (SWC), a sleep-inspired framework that scores history by salience and applies structured gist abstraction to mid-priority content, the authors evaluate four conditions over 1,935 matched questions from the LoCoMo benchmark and find gist compression beats truncation on multi-hop and single-hop factual questions but fails badly on temporal ones. They trace the failure to the gist prompt preserving relational and event structure while discarding dates and times: a one-sentence prompt modification raises temporal expression preservation roughly 20-fold (3.05% to 62.39%) while leaving entity and event preservation essentially unchanged, recovering +0.314 judge accuracy on temporal questions.
Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems
Memory systems let long-running conversational agents avoid resending the whole conversation each turn, but what they cost to serve has received little systematic benchmarking. The study compares three memory systems (Mem0, Hindsight, and Mastra Observational Memory) against a fixed-size rolling window and full-transcript resubmission, across two model backbones and conversations up to 400 turns, pairing every cost measurement with accuracy on 665 LoCoMo questions. Serving cost turns out to be unpredictable from conversation length and message size alone, with a regression that fits the reference strategies missing the memory systems by 18-69%, and the break-even point against resubmitting the full transcript ranges from tens of turns to never within 400 turns depending on system and backbone. No system wins on both cost and accuracy, which spans 21-54% across systems.
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
Reusable skills are the standard mechanism for extending large language model (LLM) agents, yet prior work reports mixed results, with some skills reducing success rates or inflating cost. The authors introduce a differential analysis framework that attributes a failure or cost regression to a specific loaded skill by comparing skill-guided runs against no-skill or semantically matched reference runs on the same task, applying it to SkillsBench and SWE-Skills-Bench to collect 307 skill-induced failures (125 functional, 182 efficiency regressions), along with a taxonomy-guided triage tool called SkillTriage. They find that functional failures rarely come from obviously irrelevant skills but from seemingly relevant ones that lead the agent to wrongly implement or omit required elements, that efficiency regressions are not explained by prompt length alone, and that excessive verification and heavy implementation pipelines are the largest sources of unnecessary mandatory work.
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Turning a research idea into a complete paper requires literature retrieval, experiment design and execution, evidence-based revision of claims, publication-ready figures, and consistency over a long generation process. The system implements this pipeline as thirteen composable skills inside an existing coding assistant, separating model-based judgment from deterministic checkable operations, specifying required evidence before results are observed, and bounding a failure mode in which repeated experiments keep rejecting the original research objective. Across eight controlled research topics it achieves 99.5% citation validity and 96.4% figure editability, raises fabrication detection from 14% for a single-pass draft to 92% with the full integrity and review stack, and averages 11.9 million tokens, $8.10, and 3.2 hours per manuscript.
LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation
Reflection — assessing trajectory progress, identifying missing evidence, and deciding whether to continue, revise, or abandon a branch — is hard for long-horizon language-model agents to learn because the decision is made locally while its utility only appears in the final outcome, leaving outcome-based reinforcement learning with sparse, delayed supervision. LoongReflect formulates reflection as a memory-control policy over a reversible trajectory tree with explicit reflect and backtrack actions, training it through two coordinated channels: a fast channel that distills globally informed reflective behavior from a privileged teacher, supervising only reflection and backtracking tokens, and a slow channel that optimizes complete trajectories with outcome-based GRPO. On multi-hop retrieval-augmented generation and mathematical reasoning benchmarks it consistently improves over outcome-only reinforcement learning and self-distillation baselines.
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
Tool-using language-model agents are typically trained and evaluated where tool calls succeed reliably, yet deployed tools fail transiently, persistently, or silently, and robust recovery may mean retrying, switching to an alternative path, or recognizing that no viable path remains. BENCH2ROBUST converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled solvability, and is used to study Bayesian Tool Memory (BTM), which provides structured runtime recovery context, alongside curriculum-controlled reinforcement learning. Across seven models from four families, injected tool failures produce a near-universal robustness gap; BTM improves robustness by up to 16.8 percentage points without retraining, reinforcement learning adds complementary recovery behavior, and combining both reaches 40.8-45.5% success under failure injection while preserving failure-free performance.
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
As AI development becomes faster and more automated, mechanistic understanding of models remains largely manual, widening the gap between what models can do and our ability to understand and control them. The proposed agentic system autonomously discovers mechanisms underlying model capabilities, supported by an interpretability-focused knowledge graph of about 13,000 papers, a 43-million-paper multidisciplinary database, and a curated library of 32 methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems it generates more valuable mechanism hypotheses and executes experiments more reliably; its reported discoveries include a safety risk in which unsafe traits transfer across modalities through apparently safe training data, a mechanistic theory of how models represent knowledge and infer beliefs, and interventions that improve performance and steer scientific foundation models toward generating DNA sequences with specified properties.
Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation
As LLM-based agents acting on user goals increasingly meet each other in strategic settings, prior theory suggests cooperation problems like the Prisoner's Dilemma become resolvable when agents know they share similar decision-making patterns. The authors build the first framework for evaluating LLM decisions when agents receive graded similarity signals about their counterpart, and complement it with a behavioral game-theoretic model of the observed reasoning. They find that models differ drastically in how they respond to similarity signals, that the dataset used to compute the signal has little effect on induced cooperation, that LLMs systematically judge other models' chain-of-thought as highly similar to their own, and that the theoretical model supports cooperative equilibria under sufficiently high similarity scores.
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS
Modernizing legacy Fortran is routine work at enormous volume, so the authors delegate it to three prompt-specialized Claude Code agent roles operating under a version-controlled specification the agents themselves authored, with humans holding a small number of approval gates and an exact verification oracle keeping the process safe. As a case study, they convert the two-electron-integral routines of GAMESS (General Atomic and Molecular Electronic Structure System) — twelve files, 56,448 lines, and 225 subroutines — from fixed-form Fortran 77 to free-form Fortran 2008, using bit-for-bit reproduction of the package's canonical test energies as the merge criterion. All twelve files pass a 51-test validation battery, and across 612 test runs the number of chemistry-relevant differences is zero.
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Enterprise agents must reason across structured APIs and document collections, but existing benchmarks test these capabilities in isolation. The benchmark comprises over 8,000 executable APIs across 62 domains with three settings of increasing difficulty — diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning under natural-language tool-use policy constraints — with correctness verified by re-executing predicted tool calls against live APIs. Using a fixed ReAct harness to isolate model capability from agent architecture, even the best frontier model reaches only 70.4% on single-hop endpoint tasks, drops to 50-51% on compositional APIs, degrades over 50% as reasoning depth grows, and falls to as low as 2.4% on unanswerable policy-constrained queries; trace analysis shows failures concentrate in language-mediated reasoning like entity disambiguation and cross-source grounding rather than tool invocation mechanics.
11 more specialized papers
- Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes Alexander Liss, Nicholas Desmond, Santiago Gil Gallego
- Harnessing agent memory to build lifelong AI partners for materials scientists Siyu Liu, Bo Hu, Beilin Ye et al.
- Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs Ruoxi Zhao, Maziar Raissi
- EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents Yuxi Qian, Yuxiang Ren
- AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search Weicheng Ye, Youran Sun, Xingyu Ren et al.
- Learning from Online User Feedback for Shopping Agents Haobo Zhang, Kelong Mao, Sulong Xu et al.
- Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents Jun He, Deying Yu
- ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models Zhou Liu, Chaoyang Han, Zewei Pan et al.
- CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations Xingyu Yan, Tingting Dai, Antonio De Domenico et al.
- Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control Josef Liyanjun Chen
- GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings Shivali Dalmia, Sumukha Thoppanahalli, Mohammadreza Sediqin et al.
Other 35
Mechanism Design for Generative Engines: From Exploitation toward Win-Win Outcomes
As generative engines make citations a key mechanism for allocating web attention and value, content providers are incentivized to game them: in repeated simulations, state-of-the-art generative engine optimization (GEO) attacks adapt to conventional defenses, producing citation-seeking rewrites that degrade document quality and introduce unsupported claims. The supplier-platform interaction is formulated as a repeated Stackelberg game with partial monitoring, and a local best-response analysis identifies when citation competition approaches an inert stationary outcome. The proposed mechanism, VCR, is based on verifiable-content rewards: rather than only penalizing suspicious rewrites, the platform credits rewrites that surface checkable factual substance. Across three benchmarks, VCR achieves the largest net defense-utility score, beating the strongest baseline by 12.1 percentage points on average and producing a win-win outcome under the paper's empirical equivalence criterion.
Principal Trait Analysis: Towards Deriving "Skills" in Human-AI Collaboration
As large language model (LLM) agents spread through the workplace, existing guidelines for effective human-AI collaboration come from either top-down theory or context-specific observations, both of which age quickly as model capabilities change. The authors propose Principal Trait Analysis (PTA), an algorithm inspired by Principal Component Analysis that uses LLM-based processing stages to derive common interaction traits from corpora of human-AI session traces, scoring each collaborator on each trait and keeping the highest-variance traits. On two collaborative coding datasets, one with students and an AI tutor and one with professional developers and a coding agent, the derived traits significantly explain collaborator behavior and help predict task outcomes, though results on whether traits generalize or remain stable over time are inconclusive.
HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging
Task vectors allow separately fine-tuned models to be merged without joint retraining, but the subset of tasks being merged varies in practice, and existing methods tune scalar coefficients for one particular subset, forcing repeated tuning and limiting merging to linear rescaling. HyperFix reframes merging over varying subsets as a combinatorial correction problem, training a lightweight hypernetwork to predict subset-conditioned nonlinear corrections in weight space; trained once on singleton, pair, and triple subsets from a task bank, it generalizes to larger subsets without per-subset optimization. A local perturbation analysis bounds the residual correction left over beyond linear merging and motivates learning it from small task updates, and experiments across several benchmarks report better results than existing merging methods at lower tuning cost.
Dion3: Full-Stack Orthogonal Updates
The Muon optimizer's cubic-time Newton-Schulz orthogonalization step is expensive, and communication overhead compounds the cost when weights are sharded across devices. Dion3 attacks this overhead at every level of the stack: a Gram Newton-Schulz algorithm reduces the FLOP cost, CuteDSL kernels exploit symmetry to accelerate computation, a megabatching strategy cuts communication, and a revised update rule orthogonalizes only a fraction of the momentum matrix's rows at each step. The result matches or improves on Muon's loss while reducing optimizer step time by up to 6x, and is released as a drop-in Muon replacement in the open-source dion package.
Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
AI tools for education are promoted as scalable fixes for access gaps, yet the infrastructure beneath them — training corpora, tokenization schemes, benchmarks, and deployment architectures — can disadvantage speakers of underrepresented languages before any model is trained. Using Bengali as a case study, the authors identify four interlocking failures: a web presence gap (under 0.5% of web content for nearly 4% of the world's population), a 67:1 English-to-Bengali training-token deficit in major multilingual corpora, a tokenization penalty from Bengali's alphasyllabary script that compounds the data deficit, and connectivity exclusion with rural internet penetration at 36.5% versus 71.4% in urban areas. They argue dataset scarcity should be treated as a structural barrier rather than an isolated technical limitation, and that offline-first design is an equity-oriented infrastructure strategy.
30 more specialized papers
- A Conceptual Framework for Refining Influence Knowledge from Simulation Evidence in Cyber-Physical Systems Barbara da Silva Oliveira (UniCA, Laboratoire I3S - COMRED, KAIROS) et al.
- Identity from the Outside: A Conceptual Framework and Research Program for AI Personality Clones Luc E. Brunet
- The Edge-based Contiguous p-median Problem with Connections to Logistics Districting Zeyad Kassem, Adolfo R. Escobedo
- VQ-bench: A Composable Vector Quantization Framework Ashwin Padaki, Amir Ingber, Edo Liberty
- Adaptive Hybrid Particle Swarm Optimization with Gradient Descent Aryan Gurudeo
- Basin: Efficient and Extensible Numerical Optimization in Rust Johan Larsson
- Socioduality: A Relational Process Framework for Human-AI Interaction Mehmed Zahid \c{C}\"ogenli
- AutoGrable: What Is a Good Graph for a Table? Tamara Cucumides, Floris Geerts
- Dual-Primal Graph VAEs for Noisy Label Aggregation Patrick Stinson, Nikolaus Kriegeskorte
- Defending against Model Extraction for GNNs with Model Reprogramming Yan Wen, Zhenyi Wang, Heng Huang
- RelShap: Relationally Consistent Shapley Explanations Seungeun Lee, Joao Fonseca, Julia Stoyanovich
- A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era Dalton Ross Smith, Wilburn Whittington, Alejandro Martinez et al.
- Sparse and robust geometric twin support vector machine via asymmetric RoBoSS loss function Kai Qi, Xinji Huang, Hongchun Wang
- Consolidator: Learning Persistent Routed Memory Across Context Boundaries Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung
- Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing Ziqiang Li, Yun Liu, Gouhei Tanaka
- High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions Hongyan Wang, Jiayu Huang, Haotian Zheng et al.
- A 12-CNOT Double Qubit Excitation Gate Irfansha Shaik
- MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning Shiji Zhou, Kunlin Lyu, Lei Zhang et al.
- HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry Haoran Pei, Zhao Su, Zetao Lin et al.
- CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation Xue Yang, Rigui Zhou, ShiZheng Jia et al.
- A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression Wouter W. L. Nuijten, Esther G. van Pelt, Albert Podusenko et al.
- TESLA: Taylor Expansion of Sinusoidal Learnable Activations Daehwa Ko, Jaehyeon Kim, Seunghyun Ham et al.
- Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision Shaojie Zhang, Ke Chen
- Towards Truly Unsupervised Evaluation of Feature Selection Hafiz Saud Arshad, Muhammad Rajabinasab, Arthur Zimek
- Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion David Bechtoldt, Sidney Bender
- Confidence Calibration of Deep Learning Systems Coby Penso
- Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning Mirko Konstantin, Stefan Zachow, Anirban Mukhopadhyay
- Structuring the Space of Perspectives Agnese Daffara, Sebastian Pad\'o, Tanise Ceron
- ADEPT: A Unified Framework for Deep Learning Test Adequacy Yidi Kao, Shawn Burnham, Tommi Rose Fahy et al.
- HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks Zhao Su, Yuxin Xia, Haoran Li et al.
Theory 33
A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph
Conway's 99-graph problem asks whether a strongly regular graph with parameters (99, 14, 1, 2) exists, and here an autonomous AI research agent mounts a systematic, fully reproducible attack scored under a partial-credit metric. The agent exhaustively proves that no circulant graph on the integers modulo 99 can satisfy more than 68% of the required constraints, and derives a forced-structure reduction showing that the parameters force each vertex neighbourhood to be a perfect matching, collapsing the existence question to that of a certain 12-regular graph on 84 vertices encoded for constraint-programming search. Its best verified construction satisfies 69.43% of constraints, a ceiling that fourteen distinct methods failed to exceed, which the authors present as a robust frontier entangled with the open problem since any provable bound below full satisfaction would constitute a non-existence proof.
Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention
Kernel attention compresses a sequence into a fixed-dimensional sketch instead of exposing every token pair as full attention does, and this work pins down exactly when that compression becomes costly. On a minimum-inner-product task over Boolean inputs, rank-one normalized kernel attention handles all sequences of length two, but any single nonnegative kernel-attention head that succeeds on all three-token sequences needs exponentially many features, while dense softmax attention solves the same task with linear-dimensional scores. The separation survives position-dependent token maps and causal queries, and the authors additionally prove transcript lower bounds for deterministic multihead, multilayer sketch models with finite-alphabet cross-token channels.
PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition
Classical PAC-Bayes generalization bounds penalize the Kullback-Leibler (KL) divergence between posterior and prior over parameters, even though predictive risk depends only on a model's behavior, and overparameterized systems contain many parameter settings with identical behavior. Using a measurable behavior map and measure disintegration, the authors decompose the classical PAC-Bayes KL divergence exactly into a behavior-selection term and a realization-level term measuring variation within fibers of behaviorally equivalent configurations, defining Z-information as the gap between the full KL and the behavior-only complexity. They further show the behavior-selection term equals the minimum KL over all posteriors inducing the same distribution of behaviors, and that symmetries, behavior-preserving directions, and fiber geometry all arise from the same structure, identifying predictive behavior as the natural object of PAC-Bayes complexity.
Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness
Convergence results for gradient descent on neural networks typically require special initializations, extreme widths, or dataset assumptions. Here convergence is proved for general feedforward networks of arbitrary width and depth assuming only that activations are Lipschitz smooth, Lipschitz continuous, and linearly bounded (true of linear, tanh, softplus, and sigmoid units) and that the loss is Lipschitz smooth in the model outputs (true of mean-squared error). The technical core is that Lipschitz properties of activations are partly preserved through repeated composition, yielding a generalized smoothness condition in which the change in gradient is bounded by the change in parameters times polynomials of the parameter norms at both endpoints, which supports a descent lemma for sufficiently small learning rates. Bounding how fast parameter norms grow then gives convergence of the minimum squared gradient norm to zero at rate O(1/T^{1/L}) over T iterations for an L-layer network.
A Quantum/Classical Example Oracle Separation for Making Things Up
A long-standing open question in the Probably Approximately Correct (PAC) learning framework is whether quantum examples confer a genuine advantage over classical examples when both learners are allowed quantum computation. The main result constructs an oracle relative to which certain distributions can be efficiently generated by a quantum learner with access to quantum examples but not by a quantum learner restricted to classical examples. This makes progress toward answering the separation question in the affirmative.
Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads
Multiplicative dual-encoder networks, which score an input pair by the inner product of separate encodings, appear across operator learning, retrieval, and contrastive vision-language models, yet lack a unified theory for choosing the number of interaction modes, normalization, or when to avoid the architecture. The authors introduce the class of functions of low interaction rank, showing approximation error splits into spectral truncation and encoder-realization terms, sample complexity scales with the sum of encoder complexities, and normalization acts as gauge fixing, with whitening pinning down interaction modes up to permutation and sign. Experiments on synthetic kernels, operator learning, and CLIP validate the predicted spectral scaling and show that independently trained CLIP models differ by a single rotation which whitening removes, exposing interpretable concept axes.
Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity
Prior results tie the number of chain-of-thought steps a bounded-depth Transformer takes to circuit complexity classes, but explicit depth-bounded constructions realizing those bounds have been missing. The authors give chain-of-thought realizations of depth-first search and Dijkstra's algorithm using unique hard-attention decoders of at most two layers, then reuse them to compute the Strahler number of an n-vertex tree in 2n-1 steps and its width in n-1 steps. Since computing the Strahler number of a binary tree is NC1-complete and the constructions need no layer normalization or positional encodings, this provides a non-trivial witness for the linear-step regime of the chain-of-thought hierarchy, with independent constructions also given on the Dyck-path representation of trees.
Disentangling the Expressivity of RoPE
Two competing explanations exist for why rotary position embeddings (RoPE) work: expressivity studies link periodic position information to modular predicates, while mechanistic studies emphasize positional anchors and local offsets. Formalizing both for finite-precision soft-attention transformers, the authors show that if every rotary component is periodic, RoPE transformers recognize exactly the languages definable in past temporal logic with modular predicates, whereas conventional RoPE, whose rotations never repeat, instead yields a precision-dependent bounded simulation of fixed-offset look-back operators. Controlled experiments match this separation: constructed periodic schedules length-generalize on modular languages, while conventional RoPE behaves like a bounded locality bias and can impair tasks requiring position-invariant access to distant context.
Reducing Symmetry Increase in Equivariant Neural Networks
Equivariant neural networks lose expressivity on symmetric inputs: passing a symmetric input through an equivariant map can only increase its symmetry, making outputs invariant to transformations beyond those of the input itself. The work gives a rigorous characterization of this phenomenon, proving that for any feature space and input symmetry group the increased symmetry admits an infimum determined by the feature space's structure, providing a computable algorithm to derive that infimum, and distilling feature-design guidelines that provably reduce harmful symmetry increase for most equivariant maps under standard regularity assumptions. Visualizations and experiments on synthetic data and the QM9 molecular dataset validate the theoretical predictions.
NAE: Normalizing AutoEncoder
Normalizing flows with approximate inverses, spanning both full-dimensional and bottleneck settings, are grouped here under the term flow autoencoders, and a theoretical analysis of their training dynamics proves that the loss used by existing approaches is suboptimal: encoder and decoder surrogates must be optimized in alignment with the reconstruction loss. Guided by this, the proposed Normalizing Autoencoder (NAE) uses a conditional loss that aligns the surrogate loss gradient with the reconstruction gradient. Experiments on molecule generation, tabular data, and image benchmarks show state-of-the-art generative performance, supporting loss alignment as a key design principle for this model family.
23 more specialized papers
- WavePhaseNet: A DFT-Based Method for Constructing Semantic Conceptual Hierarchy Structures (SCHS) Kiyotaka Kasubuchi, Kazuo Fukiya
- Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning Suyash Mishra
- Every pooling rule has its world: matching probability combination rules to situations and stakes Tanel Tammet, Priit J\"arv, Dirk Draheim
- Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction Yi Liu
- Spectral graph clustering with inhomogeneous latent geometry Konstantin Avrachenkov, Lucas S. Sibemberg, Alexander Van Werde
- RevCRN: Reversible Analog Computation using Chemical Reaction Networks Saptarshi Biswas, James I. Lathrop, Rana D. Parshad
- Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints Zhen Xu
- Strengthening Full Justified Representation: Efficient Verification and Computation Nicholas Teh
- On Weak Bisimilarities in CCSK Baptiste Vall\'ee, Ivan Lanese
- Fine-Tuning Generative Models for Extreme Events via CVaR-Penalized Wasserstein Gradient Flows Thejani Gamage, Hyemin Gu, Zhizhen Zhang et al.
- Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan et al.
- A Local Sinkhorn Framework for Conditional Distribution Reconstruction of Multidimensional Random Fields Mingtao Xia, Qijing Shen
- Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning Tieliang Gong, Zhongbo Zhang, Wen Wen et al.
- Proportional Analogies on Probability Distributions via Bayesian Updating Pierre-Alexandre Murena
- Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp Wenzhi Gao, Zhaonan Qu, Yinyu Ye et al.
- Kernel Methods for Learning Operators with Multiple Inputs and Outputs Adrien Weihs, Chunyang Liao, Jingmin Sun et al.
- DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks Andy Wang, Charlton Shih, William Chang
- Causal inference for group-contaminated structured outcomes: observable quotients, lossless reduction and exact randomization inference Usef Faghihi, Amir Saki
- Adaptive Bregman Proximal Stochastic Gradient with a Stabilized Barzilai--Borwein Step Size Chenhan Jin, Shengze Xu, Binghui Xie et al.
- Direct Acceleration of Stochastic Root-Finding Without Variance Reduction and Regularization TaeHo Yoon, Nicolas Loizou
- The Advective Fisher-Rao Geometry of Deterministic Measure Transport Benjamin Gess, Johannes M\"uller
- Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids: From Robust Offline Optimization to Full-Bandit Learning Vaneet Aggarwal
- An Efficient Near-Optimal Algorithm for Adversarial $m$-Set Bandits Francesco Bacchiocchi, Tommaso Cesari, Roberto Colomboni
Safety & Alignment 25
Forecasting Side Effects of Activation Steering
Activation steering modifies a language model by adding a learned direction to its hidden activations, but it often produces unintended side effects on other behaviors, making safe deployment difficult. The authors construct a cross-effect matrix over a taxonomy of 67 behaviors across three open-weight models, finding that side effects are common, structured, and often asymmetric in ways existing similarity-based heuristics cannot explain. Despite this complexity, the effects prove largely predictable before steering is applied: their magnitude depends primarily on the target behavior, and their direction can be forecast from the model's unsteered representations substantially better than simple baselines, enabling proactive auditing of steering interventions.
Agent Safety Should Be a Runtime Contract
The dominant paradigm treats AI safety as something instilled during training, which this position paper argues is structurally insufficient for autonomous agents that execute code, mutate files, send messages, and modify databases; safety should instead be a runtime contract enforced by the agent harness. The contract has a preventive face — sandboxes, permission gates, output filters, and trajectory monitors that block dangerous actions — and an evidential face that gates task submission on verifiable proof such as test runs, log captures, file diffs, and citation grounding. The position is supported by a survey of 52 documented agent and LLM safety incidents, a false-completion audit, a trajectory-schema audit of 12 public agent systems, and a title-level audit of 28,560 papers from NeurIPS, ICML, and ICLR 2023-2025 showing a pooled 8-12x publication imbalance favoring training-time over deployment-time safety; the authors formalize an Agent Trajectory Schema and Evidence Chain and outline a research agenda.
Backdoor Decontamination Dynamics in LLM Agents
Open-weight language-model agents can carry backdoors installed during fine-tuning that stay invisible if their trigger never appears at test time, and a defender who does not know the trigger cannot unlearn it directly. One proposed remedy is defensive poisoning: deliberately install a backdoor you do know, then unlearn it and hope the unknown one is removed as a side effect. Using AgentDyn, a tool-calling agent testbed that varies trigger, response, teacher, and fine-tuning method independently, 115 experiments find defensive poisoning alone erases roughly 56% of original backdoors while the subsequent unlearning step eliminates nearly all survivors, and malicious backdoors never persist when their trigger is of the same general type as the defensive one. Co-installing up to four backdoors raises resistance (about 36% erased), yet decontaminating a single known co-resident clears 52 of 60 others, and inspection of internals with J-lens shows benign responses are restored while traces of trigger awareness remain in intermediate layers.
Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment
Group alignment adapts a language model toward a demographic group's opinions, values, and preferences, while sycophancy, a known by-product of alignment, makes models over-agree with users regardless of the facts; existing group alignment work measures only opinion fit and ignores the sycophancy it induces. Group Alignment-induced Sycophancy (GAS) evaluates both sides together across three alignment methods, four models, and 13 demographic groups, tracking the intended gain in opinion alignment alongside the unintended shift in sycophantic behavior. Both effects are non-uniform: under an identical budget some groups gain far more opinion alignment than others, and the induced sycophancy forms a group-specific multi-dimensional profile rather than a single shift. The authors conclude that group alignment should be reported as a two-sided profile accounting for per-group differences rather than one fit score.
EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval
Safety alignment in large language models is often treated as a property spread across the whole network, yet its brittleness suggests refusal behavior may be concentrated in relatively few parameters. By transplanting weights from aligned models into matched unaligned base models at several granularities—attention weights, multilayer perceptron (MLP) weights, contiguous layer regions, and individual MLP blocks—across two open-weight model families and four safety benchmarks, the study finds that refusal transfer is dominated by MLP weights, which recover at least 2.7 times more refusal of malicious prompts than attention weights, and is concentrated mid-network, with the block spanning layers 8-11 selected first in all six greedy searches. Composition is non-additive: adding more aligned blocks can reduce refusal performance, and selective subsets can beat full MLP transplantation, indicating that alignment is both localized and interaction-sensitive.
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
As large language models (LLMs) increasingly debate, advise, and collaborate, resistance to harmful persuasion becomes a reliability requirement — yet a single targeted persuasive argument, even a factually false one, can collapse a model's accuracy to near zero. An adversarial reinforcement learning framework trains persuader agents to flip a target model's answer in a single interaction, raising persuasion success from roughly 24% to over 93% against the training-time target, with learned strategies transferring to unseen models (83% on Qwen-14B, 79% on Llama-3.1-8B, and 25% on GPT-4o-mini, rising to 38% with a curriculum that bootstraps on more persuadable models first). The optimized persuaders increasingly rely on credibility-based tactics such as fabricated citations and false authoritative evidence, leading the authors to position persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making.
LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection
Multimodal large reasoning models (MLRMs) trained with reinforcement learning can leak supposedly unlearned sensitive facts inside their chain-of-thought traces even when the final answer is sanitized, a vulnerability much more pronounced than in non-reasoning base models. The authors find that sensitive content in these models carries a distinctive token-level entropy signature, and build LEMUR, a training-free inference-time unlearning framework that uses entropy dynamics to detect when sensitive reasoning begins and then redirects the trajectory by injecting sanitized, probability-weighted embeddings re-grounded in the input image. Across diverse models, the method suppresses both reasoning-trace and answer leakage better than existing unlearning approaches while preserving utility on non-sensitive content and output fluency.
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning
Aligned large language models can flip their safety decisions on an identical request depending on the persona traits assigned in the system prompt, a failure mode the authors call trait-induced safety variation and measure with two refusal-based metrics. A representation-level analysis shows traits perturb the model's safety representations within a low-dimensional subspace, motivating Trait-Invariant Safety Tuning (TIST), a self-distillation framework that aligns trait-conditioned behavior with the model's no-trait behavior, and an instantiation called Trait-Subspace Neutralization (TraSN) that enforces invariance only within that subspace. Experiments show TraSN stabilizes safety behavior across traits and strengthens refusal of harmful requests while preserving general capability.
Locating and Controlling Implicit Personalization in Large Language Models
Large language models shift their outputs in response to implicit demographic cues even when users never state an identity, but how this behavior connects to internal activations has been unclear. Using matched cued and neutral conversations across five models, the authors find a localized activation signal that tracks recommendation changes with correlations up to r=0.87, and show that removing the signal for a given cue suppresses its influence, often more effectively than prompting the model to ignore demographics, while largely preserving benchmark performance. Selectively removing one demographic dimension while leaving co-present ones intact, however, remains highly model- and attribute-specific.
Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
Regulators increasingly require disclosure of AI involvement in public communication, but whether such disclosures actually blunt persuasion has been unclear. In a preregistered experiment, 1,500 UK adults held a short conversation with an identical persuasive chatbot about one of 60 policy issues, randomized to receive no disclosure, a prominent AI-identity disclosure, or that disclosure plus the chatbot's persuasive intent and instructions. Identity disclosure alone was practically equivalent to no disclosure (13.1 versus 12.6 points of attitude shift on a 100-point scale), while adding intent disclosure roughly halved persuasion to 6.3 points and made participants view the campaign as less acceptable — suggesting regulation of persuasive AI must address what a system is trying to do, not just what it is.
How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment
State-aligned distortion is documented in China-origin text-based large language models (LLMs), but its multimodal form had not been systematically examined. The study benchmarks nine vision-language models (VLMs) — seven China-origin, two not — on 200 entries spanning ten politically sensitive topics plus a visual-abstraction probe, across four elicitation paradigms and two prompt languages for 21,708 trials, each audited on six dimensions by two frontier LLM judges validated against human experts. Chinese-language prompting roughly triples the odds of state-aligned framing in every model; China-origin models reframe 1.6-3.2x more than non-China models; the effect is gated by recognition of the depicted subject rather than pixel detail; and across four Qwen generations, framing rises while explicit refusal falls, meaning censorship migrates from a visible act users can notice to invisible fluent reframing.
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
Large language model (LLM) agents that call external tools are vulnerable to indirect prompt injections hidden in environment state, but existing security studies rely on hand-built environments and predefined injection locations. ToolHazard is a scalable framework that synthesizes executable, stateful adversarial environments using an Environment Simulator, an Attacker Agent, and a User Simulator, which together discover viable injection points, generate environment-specific payloads, and construct long-horizon tasks. The resulting ToolHazard-Bench reveals substantial agent vulnerabilities and shows that injection timing and placement affect attack success, while alignment data generated by the framework improves security on both ToolHazard-Bench and AgentDojo without degrading benign task performance.
Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed
Small language models (SLMs) are built either by training compact models from scratch or by compressing larger pre-trained models, but how these paths affect trustworthiness is underexplored. The study evaluates SLM trustworthiness across fairness, robustness, privacy, and ethics, finding that quantization preserves trustworthiness significantly better than pruning, and that quantizing a reliable large model produces SLMs with better trustworthiness and adaptability than small models trained from scratch. Knowledge distillation from trustworthy teacher models further enhances reliability, offering practical guidance for deploying trustworthy small models.
No One to Blame: A Framework of Constitutive AI Unaccountability
Most research treats AI accountability gaps as barriers that better standards, transparency, or institutional reform can overcome; the authors argue instead that some configurations of actors, systems, and institutions make accountability conceptually unachievable, a condition they call constitutive AI unaccountability. Through a concept-centric literature analysis, a secondary analysis of 27 expert interviews with technical, legal, and sociotechnical AI professionals, and an application to the open-source agentic system OpenClaw, they identify nine categories and 20 themes spanning structural, technological, and normative clusters linked by eight directed interdependencies. Their 20-question diagnostic instrument detected 17 of the 20 conditions in OpenClaw, including an inverted-anthropomorphism configuration in which the AI agent was the only identifiable actor.
Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents
LLM agents that rely on third-party skills expose two control points to untrusted publishers: natural-language descriptions used for skill selection and instruction bodies used for planning. Convergent Detour Hijacking (CDH) is a text-only, runtime-independent attack that couples these stages under shared semantic cover — a description establishes relevance during selection, then an aligned body fabricates plausible dependencies that recruit unnecessary benign skills into a bounded detour before re-entering the original route, so the task still completes. Across multiple LLM backends and 491 held-out tasks, the attacker-controlled coordinator is selected in 80.02% of tasks on DeepSeek-V4-Pro, inflating token consumption by 66.91% and execution time by 92.45% while task completion stays comparable — showing that correct outcomes do not guarantee trajectory integrity or cost safety.
10 more specialized papers
- The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification Yoshinori Watanabe
- Variable Selection in the Context of AI Fairness Ivan Luciano Danesi, Chiara Frigerio, Fabio Maccaferri et al.
- Governing Agentic AI in FinTech Henry Han
- Generative Learning for Quantum Measurement Design Jun Dai, Olivier Nahman-L\'{e}vesque, Guillaume Rabusseau et al.
- Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits Hangqi Ren, Junyi Liao
- Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark Soumya Mazumdar, Vineet Kumar Rakesh, Tapas Samanta
- Robust Ambiguity Detection (RAD) From Model- and Feature-Space Consistency Manya Singh, Mark T. Keane, Arjun Pakrashi
- Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study Simone Mungari
- Clustered Randomized Smoothing for Stochastic Prediction Functions Eduardo Figueiredo, Frederik Mathiesen, Julian Schumann et al.
- Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers \'I\~nigo de Troya, Maurus Enbergs, Neelke Doorn et al.
Vision 24
LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration
Sampling video diffusion transformers is dominated by quadratic self-attention over long 3D token sequences, and existing training-free sparse-attention methods pursue aggressive sparsity where extra speedup costs disproportionate fidelity. This method targets the opposite regime: it fixes a 99% retained-attention-mass threshold, measures exact block attention masses at one early dense denoising step, keeps for each head and query block the smallest set of key/value blocks meeting the threshold, and freezes those indices for all remaining steps — exploiting the finding that roughly 40% of block interactions can be dropped while the high-mass support stays stable across steps. On Wan2.1-1.3B it delivers a 1.36x speedup with only a 0.06-point drop on the VBench quality benchmark, and combined with feature caching reaches 3.2x on HunyuanVideo at a 0.02-point drop, versus 0.32 points for the strongest sparse baseline at comparable speed.
How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging
Deploying unsupervised domain adaptation (UDA) in clinics requires choosing both an algorithm and a specific trained model, but the deployment domain is unlabeled, so candidate models cannot be evaluated on it directly. The study evaluates the complete pipeline — adaptation together with label-free model selection — across eleven clinically relevant cross-domain scenarios from nine medical imaging datasets, ten UDA algorithms, and 13 label-free selection methods, covering over 80,000 trained models. A capable adapted model usually exists, but no evaluated selection method finds it reliably, leaving a large structural gap to the best available model; ensembling and a small target-labeling budget narrow the gap without closing it, suggesting the under-explored selection step is what keeps UDA from clinical use.
22 more specialized papers
- Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning Shibo Gao, Peipei Yang, Xu-Yao Zhang et al.
- SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation Dongsu Song, DaeYun GO, Boseung Seo et al.
- CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification Gawon Lim
- Uncertainty-Aware Compositional Localization and Placement Assessment of Catheters and Tubes in Chest X-Rays Harshil Lodhiya
- Gloss-Free Representation Learning for Cross-Dataset Sign Spotting O\u{g}uz Akif T\"ufekcio\u{g}lu, Ezgi Ekin, Mustafa Kaan \c{C}evik et al.
- Gaze Target Estimation Anywhere with Concepts Xu Cao, Houze Yang, Vipin Gunda et al.
- Click2Poly: A VLM for vector mapping buildings and walls Nicolas Girard, Jawher Ben Abdallah, Arno Gobbin et al.
- Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment Weize Cai, Yongqi Dong, Zhida Shao et al.
- From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection Zepeng Wang, Jiagao Hu, Fuhao Li et al.
- CAM-Guided Saliency Cutout and Image-Based Malware Classification Yasaman Ebrahimi, Martin Jurecek, Mark Stamp
- Robustness of AI-Art Detectors under Generator Shift Shivank Singh Thakur, Meien Li, Mark Stamp
- Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang
- Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation Yuanmin Huang, Chen Chen, Geng Hong et al.
- Can Vision Models Read the Radar Display? On the Feasibility of Radar Imagery for Air Traffic Complexity Estimation Hyewook Kim, Byul Kang, Seokbin Yoon et al.
- Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks Yaohua Liu, Yifan Guo, Jiaxin Gao
- Draw This First Dazhi Zhong, Rowan Bradbury, Grant Davis
- Autonomous Telerehabilitation via Skeletal Motion Prediction and Joint-Level Performance Assessment Lara Pereira, Jo\~ao Ruivo Paulo, Pedro Santos et al.
- HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation Ruochen Li, Shuang Chen, Wenke E et al.
- M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation Jing Zhu, Ye Wang, Fumin Wang
- HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression Yuefeng Zhang
- A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery Rafi Ibn Sultan, Chengyin Li, Yiannos Demetriou et al.
- Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini et al.
Reinforcement Learning 19
PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR
Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories, and existing budget allocators assign compute using pointwise notions of prompt or rollout difficulty. The authors show this mismatches the actual gradient structure: the leave-one-out group-relative score gradient is a second-order U-statistic over pairs of rollouts, so completing one rollout reveals contrast with every other. Their method, PAIR (Pairwise-Aware Inclusion Reweighting), uses a prefix-only predictor to set continuation probabilities under a token budget and inverse-weights each observed pair term by its joint inclusion probability, giving a design-unbiased gradient estimator. In compute-matched runs on Qwen3-1.7B and 4B, PAIR improves accuracy by 1.2 and 1.4 points over the strongest pointwise allocator while generating about half the tokens of full-group GRPO.
TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs
Standard evaluation metrics for offline reinforcement learning (RL) in medicine, such as mean squared error (MSE) and Fitted Q-Evaluation (FQE), measure only behavioral imitation and so cannot detect 'toxic mimicry' — agents replicating harmful patterns like treatment withdrawal during comfort-care transitions. The Counterfactual Clinical Audit (CCA) framework stress-tests agents trained on the MIMIC-III intensive-care database using physiological perturbations anchored in Surviving Sepsis Campaign (SSC) guidelines. The audit shows that a Medical Decision Transformer paradoxically reduces vasopressor dosage as lactate escalates, contradicting resuscitation guidelines, while a Historical Causal Transformer using causal action shielding, propensity-based weighting, and Conservative Q-Learning maintains physiologically consistent responses — evidence that statistical fit and clinical safety can systematically diverge.
Let it Cook: Learning to Wait in Sequential Decision Making
Agents in sequential decision-making problems usually sense and act at every timestep, but many tasks contain stretches, such as waiting for coffee to brew, that are served just as well by letting the environment evolve unmonitored. The authors formalize learning to wait as minimizing how often the agent senses and decides without sacrificing task performance, and train a waiting policy that decides where and for how long to pause, committing to a fixed number of timesteps without sensing, using reinforcement learning with lexicographically ordered objectives. Across four discrete-state household tasks and three continuous-state environments the method learns waiting behavior and can adapt pre-trained policies to wait where appropriate; how much waiting each task tolerates differs, but the learned solutions sometimes wait for over half the task duration.
Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning
In single-agent reinforcement learning, successor features with generalized policy improvement let a library of learned policies be recombined for any new objective with a guarantee of not underperforming the library, but the common multi-agent practice of letting each agent recombine its own library independently loses that guarantee. The authors prove independent composition can yield joint behavior strictly worse than every library policy, because recombining teammates changes each agent's effective environment, and that the only unconditionally safe fixed rule, synchronized composition, cannot assign different goals to different agents. Their proposed MA-USFA is a hierarchical method pairing universal successor feature approximators conditioned on teammates' objectives with an upper-level composer that picks each agent's library entry and supplies a cross-agent value correction, trained once over an objective distribution and deployed without per-task adaptation.
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL
Reinforcement learning against rubrics graded by a large language model judge is prone to reward hacking, since the rubric is only a proxy for quality. Training Qwen3-8B with Group Relative Policy Optimization (GRPO) on medical and science rubrics, the authors show the training judge's score keeps rising while a stronger gold judge's score peaks and then falls, by 3 points on HealthBench-Hard and 22 points on ResearchQA, a divergence they argue cannot be explained by fixed judge bias. Their fix, Rubric Dropout, randomly drops a subset of rubric criteria before computing each reward (shared within a rollout group so advantages stay comparable), which raises out-of-distribution gold scores at every checkpoint, reduces measured hacking, and works across a broad 30-50% dropout range, while reweighting criteria by usefulness performs worse than no intervention.
GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs
On-policy rollout methods like Group Relative Policy Optimization (GRPO) for post-training large language models suffer from instability, cross-task capability degradation, and response-length inflation. The authors introduce Principal-Subspace Overlap, a dimension-corrected measure of how individual rollout updates align with the dominant singular subspaces of pretrained weights, and find that transient overlap spikes often precede performance drops. Their method, Geometrically Constrained Policy Optimization (GCPO), applies hard bilateral orthogonal projections that confine updates to the complementary subspaces, and across math, code, and tool-use tasks on Qwen3-8B and GLM4-9B it outperforms GRPO, DAPO, and GSPO by up to 2.37 points over the strongest baseline while preserving general capabilities and eliminating length inflation.
Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models
Object-centric world models, which decompose a scene into slots bound to individual objects, have been proposed as a more sample-efficient and generalizable basis for planning, but prior work took the slot encoder as given and only evaluated in-distribution. This study systematically varies object-centric representation quality and tests generalization under distribution shift in visual model-predictive control, comparing against scene-centric models. It finds that planning success correlates with unsupervised slot-quality metrics (saturating at high quality), that well-bound slots make previously used proprioception inputs and masking biases unnecessary, and that while slot-based models are more robust under unseen shifts than an end-to-end scene-centric baseline, DINO-WM built on frozen pretrained features remains competitive, suggesting pretrained features are a key driver of robustness.
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Multi-agent reinforcement learning for human-AI interaction typically trains a policy against a single large language model that simulates users, and the authors show this fails to generalize: because the simulator is mode-collapsed, the policy overfits to strategies that exploit its dominant behavior mode and transfers poorly to unseen simulators and real users. They formalize this "simulator collapse" and propose two remedies — Verbalized Sampling, which broadens simulator behavior at inference time by sampling from a verbalized response distribution, and Co-Training, which optimizes the policy against a population of trainable simulators. On Persuasion for Good, tau^2-bench, and CooperBench, Verbalized Sampling improves held-out success by up to 9% over single-simulator RL and Co-Training by up to 14%, with similar gains in a human study; the authors also release SCOPE, an open-source framework for population co-training.
A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions
Non-experts struggle to build reward functions that adhere to a given preference ordering over trajectories, so the authors formalize a three-step process that turns a natural-language task description into a linear reward function. The steps are: a guided workflow distilling the task into fundamental objectives with measurable outcome variables, a reduction of reward-term selection to minimum-cost partial cover on a causal directed acyclic graph solved in polynomial time via max-flow, and weight fitting framed as a convex feasibility problem iteratively narrowed by preference queries using separation oracle methods. The method is claimed to be the first to maintain a deterministically conflict-free feasible weight region, narrowed to a desired tolerance with O(n log kappa) preference queries.
10 more specialized papers
- Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version) Jack Mirenzi, Henny Admoni
- Reinforcement Learning based DBMS Buffer Pool Auto-Tuning for Optimal Memory Utilization Yifan Wang, Patrick Royer, Rapha\"el F\'eraud et al.
- Dueling Deep Q-Learning for Intrusion Detection Logan Luna (Georgia Institute of Technology), Matthew P. Berkowitz (Embry-Riddle Aeronautical University), Laxima Niure Kandel (Embry-Riddle Aeronautical University) et al.
- Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings Tran Le Vu
- Dynamics Models for Offline Hyperparameter Selection in Real-World RL Jordan Coblin, Han Wang, Martha White et al.
- Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving Aditya Humnabadkar, Huaizhong Zhang, Ardhendu Behera
- When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits Sang Su Lee, Vineeth Loganathan, Shishir Dash et al.
- GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation Ofir Ben Shoham, Shrutendra Harsola, Vignesh Subrahmaniam et al.
- Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation Md Yassir Mottalib, Md Yousuf, Eklachur Rahman Bhuiyan et al.
- Redistribution-based Cost Inference Improves Sparse Safe Offline RL Ebenezer Gelo (University of the Witwatersrand), Geraud Nangue Tasse (University of the Witwatersrand), Steven James (University of the Witwatersrand) et al.
Multimodal 18
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval
Retrieval and classification across text, images, video, and audio has traditionally used dual-encoder models trained with contrastive learning, and the March 2026 release of Gemini Embedding 2, which maps all of those modalities plus documents into one shared space, changes the competitive picture. Because frontier large language models also show strong visual understanding, the question is whether they can act as effective zero-shot rankers instead. A direct comparison on Flickr30k with hard negatives finds GPT-4.1 and Claude Sonnet 4.6 performing on par with the native multimodal embedding model, though once embeddings are precomputed the embedding approach is the better fit for low-latency applications.
CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models
Computed tomography (CT) matters clinically not only for describing current disease but for comparing serial scans to track how disease evolves, which underpins treatment-response assessment and recurrence detection, yet medical foundation models mostly handle one study at a time. CT-DeltaBench targets longitudinal difference reporting, where a model receives two temporally separated scans from the same patient and writes a report describing the interval changes, using patient-level splits to prevent leakage, change-aware metrics that go beyond surface text similarity, and independent physician validation of the synthesized reference reports and event-extraction pipeline. The authors compare direct reasoning over the paired scans against a two-stage pipeline that produces single-timepoint reports and then diffs the text, and release DeltaMed, a baseline trained on the benchmark for direct paired-CT difference reporting.
Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting
Multimodal large language models (MLLMs) typically process videos through sparse uniform frame sampling to control token costs, which can discard the critical transitions needed to reason about object movement, collisions, and causal interactions. Motion-as-Prompt recovers dense point trajectories, selects motion-informative frames, and marks the trajectories accumulated between consecutive sampled frames directly onto the visual inputs, making otherwise hidden displacement and direction changes visible to frozen models. Without any training or architectural modification, it improves average motion-reasoning accuracy for GPT-5.5 by 4.2% on CLEVRER and 8.9% on Something-Something-v2, without degrading non-motion understanding.
LookBack: Where and How to Score LVLM Responses via Visual Reference Usage
Large Vision-Language Models (LVLMs) hallucinate not just at the text level but against the image itself, producing fluent responses ungrounded in what they see. The authors' diagnostics show that confidence-based scoring metrics adopted from large language models barely change when the input image is removed, meaning they capture textual plausibility rather than agreement with the image. LookBack is a training-free response scoring method that augments token likelihood with a lightweight visual lookback score measuring how strongly each response token refers to image tokens. Across four benchmarks and three models, it consistently improves Best-of-N response selection over existing baselines with negligible overhead.
Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models
Unified multimodal models (UMMs) aim to combine visual generation and understanding in a single parameter space, yet current evaluation protocols test the two capabilities separately, leaving no system-level assessment. Self-Generative-Understanding (SGU) is an annotation-free framework that closes a semantic loop: the model first describes an image, then reconstructs a visual context from its own description, and finally reasons over that self-generated output, yielding an integrated performance score at zero annotation cost. Experiments show that even high-performing UMMs often struggle to reason over their own generated contexts, revealing limitations that separate evaluations of understanding or generation fail to capture.
SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward
Vision-language models remain weak at spatial reasoning, and reinforcement learning that rewards only verifiable final answers gives no signal about which intermediate step was responsible. SCOUT pairs a structured chain-of-thought format that explicitly writes out 3D perception of the environment, including the depth cues other structured-reasoning approaches skip, with a reinforcement learning algorithm using multi-objective process rewards and a tailored advantage estimator that assigns credit to distinct segments of the reasoning trajectory; training uses SCOUT-24k, a synthesized dataset of structured spatial reasoning traces. The 3B model improves on its baseline by 16.85% on general spatial benchmarks and 6.3% on complex spatial reasoning tasks, and the 7B model reports a 4.28% margin over GPT-4o. Although trained only on single images, the 7B model transfers to multi-image and video inputs.
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Multimodal large language models (MLLMs) are increasingly used in scientific writing workspaces that convert diagrams directly into LaTeX TikZ code, but their diagram abilities lacked systematic evaluation. The benchmark provides 3.7k curated scientific diagrams and 18.3k human-validated questions across six domains, testing three tasks — diagram-to-code parsing, diagram-to-code editing, and diagram question answering — each also in an agentic setting. Evaluating 12 models shows that diagram-to-code tasks are markedly harder than question answering: models reason well over diagrams but struggle to parse and edit them, and while agentic settings improve parsing and editing for most models, they degrade question answering for all but Claude-4.6 Opus.
AVA-Encoder: Towards Agent-Native Video Representation Learning
Creative agents lack a structured video representation that both stays faithful to film content and can be directly queried and edited during agentic reasoning. The Agentic Video Auto-Encoder addresses this by encoding a video into a knowledge graph whose nodes hold structured text linked to generated image, audio, and video assets, then reconstructing the video from that graph; differences between the original and reconstruction drive a textual-gradient loop that rewrites the encoding policy in natural language and can optionally refine individual graph representations at test time. The authors report a 20.7 percentage-point improvement over the strongest external baseline, and their pseudo-trained shot-level encoding policy outperforms a carefully human-tuned one while using 74.3% fewer system-prompt tokens, alongside released code, a video reconstruction benchmark, and a dataset of film knowledge-graph representations.
10 more specialized papers
- Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation Md Maklachur Rahman, Tracy Hammond
- ODE-Based Transformer Decoders for Iterative Sign Language Translation Tu\u{g}\c{c}e K{\i}z{\i}ltepe, Hacer Yalim Keles
- TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation Haoqi Hu, Tongji Luo, Li Zhang et al.
- IoT-Enabled Autonomous Maritime Navigation in Smart Ports: A Curriculum-Guided Shared Policy Learning Framework Yuqing Lin, Rangya Zhang, Kum Fai Yuen
- MBA: Multimodal Benchmark and Agents for Real-World Business Ideation Hojun Choi, Jaeyo Shin, Suin Lee et al.
- Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models Myeong-Ju Cho, Hye-Bin Shin, Seo-Hyun Lee et al.
- Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System Haoyu Zhang, Shuoxun Zhang, Peng Ye et al.
- JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis Ran Li, Huiguo He, Jiahuan Cao et al.
- AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention Juncheng Liao, Jinfan Lv, Guoming Wang et al.
- Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence Aman Tyagi, Hemanth Boinpally, Jonathan Chen et al.
Robotics 11
Towards the Harness of Embodied Agents
Coding agents established the harness paradigm, in which what an agent achieves depends not on the model alone but on the infrastructure around it, and the question posed here is whether that paradigm extends to embodied agents in the physical world. The harness runs an agentic loop that orchestrates robot capabilities, each wrapped as a callable tool, inheriting the core components of coding agents with modifications the physical world requires. Because the physical world does not freely grant the ability to read state or judge action outcomes, the system introduces Scene Graph as Context, a persistent symbolic representation of the world, and Evaluation as Exit Codes, which detects when an action should terminate, judges whether it succeeded, and diagnoses failures. Together these close the loop between agent and environment, letting composed tools carry long-horizon tasks to completion in real settings.
Adaptation of Generalist Robot Policies with Minimal Data
Fully autonomous robot learning remains out of reach because sparse rewards and weak zero-shot exploration make it unlikely a robot discovers successful behavior from scratch, so the authors study minimal-data adaptation: learning a new task from as little as one demonstration followed by autonomous online interaction. Their recipe, MiDAS, first anchors a pre-trained vision-language-action (VLA) policy to the target task with behavior cloning on the demonstrations, then improves it through value-based online reinforcement learning on a residual policy parameterization. Across the LIBERO and RoboCasa benchmarks, MiDAS recovers strong performance from a single demonstration and generalizes beyond demonstrated conditions, and on a bimanual YAM platform it turns a fragile single-demonstration policy into a more robust one over about six hours of online interaction — claimed as the first reliable robot policy adaptation from a single task demonstration.
Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards
End-to-end autonomous driving agents can score well on average yet still violate basic traffic rules, because they learn statistical patterns rather than the physical conditions that guarantee safety. The authors add a lightweight neuro-symbolic guard to the final command interface of an already-trained agent: each command is checked against explicit safety rules just before execution and, only when a rule is violated, replaced with the nearest safe alternative, with every intervention traceable to its triggering rule and no retraining required. Applied to the state-of-the-art TransFuser v6 driver on the Fail2Drive and Bench2Drive long-tail benchmarks, the guard improves success rate by 15% and cuts safety-critical collisions by up to 53% while preserving the original driving score.
Keep the Future, Drop the Rollout: RIFT for World Action Models
World action models (WAMs) condition robot actions on predicted futures, but generating those futures by iterative video rollout adds deployment latency, raising the question of whether the action policy needs the evolving rollout trajectory or only its final future representation. Paired closed-loop interventions on four WAMs across all 40 LIBERO tasks show policies are sensitive to future-cache values and their assigned positions, yet for the Joint and Cosmos-2 models replaying one fixed final-clean key/value cache nearly preserves behavior, with 1.7 to 1.9 cm average end-effector displacement error and 97.9% to 98.2% success, separating cache consumption from cache production. RIFT (Rollout-free Imagination via Future Tokens) exploits this by using learned anticipation tokens to build a complete future key/value cache in a single backbone pass while keeping the original future-read interface. It reaches 98.8% success on LIBERO against 98.4% to 98.6% for rollout-based baselines while cutting action-chunk latency by 68.2% to 89.1%, and scores 92.9% and 92.6% on clean and randomized RoboTwin 2.0 scenes.
Foresight Without Seeing: Latent Futures for World Action Models
World Action Models (WAMs) couple future visual prediction with robot action generation, but explicit-future variants pay heavy inference costs for iterative video denoising while direct-policy variants give the action model no access to predicted dynamics. ForeWAM bridges the two: a single Video DiT prefill over the current visual latent and stochastic future slots produces layer-wise key-value states that are reused throughout action denoising, and dynamics registers supervised by a frozen latent-action teacher push these implicit future states to capture object motion, contact changes, and task progress. Deployment requires no future video generation, and without embodied robot pretraining the method reaches 96.7% average success on LIBERO and 61.6% on LIBERO-Plus.
G0.5: One Autoregressive Stream for Robot Reasoning and Action
Most vision-language-action (VLA) models pair a pretrained vision-language model with a separately trained flow-matching action expert, reducing the language model to a context encoder rather than a decision-maker. G0.5 instead has a single transformer decoder emit reasoning and action tokens under one objective, made tractable by a learnable cross-embodiment action tokenizer, a native chain-of-thought stream interleaved with actions, and a visual memory module injecting multi-second history. Because reasoning and action share weights, prompts directly steer action granularity, task horizon, and out-of-distribution handling, and the model surpasses state-of-the-art baselines across seven regimes, including 76.7% on real-world robot fine-tuning versus 53.3% for pi-0.5 and 98.9% on LIBERO.
Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
Scaling reinforcement learning to tasks that combine locomotion and manipulation is bottlenecked by the slow, manual process of dense reward shaping. The approach instead uses sample-based model predictive control (SMPC) in simulation as an automated, rapidly tunable expert to generate large offline datasets, which solve the exploration problem well enough that an off-policy agent can be trained with purely sparse task rewards, and pairs the resulting high-level policy with a low-level dynamic stability controller. The learned policies align strictly with true task objectives and ultimately surpass their optimal-control teacher, with sim-to-real deployment demonstrated on an arm-equipped Spot quadruped and a G1 humanoid.
Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment
Learned motion planners handle complex traffic scenes better than hand-written rules but are opaque, which undermines explainability and complicates safety assurance. The proposed hybrid architecture keeps both properties: a deep neural network interprets the scene and proposes a driving behaviour, while an optimization-based supervision layer validates that proposal and enforces explicit drivability and safety constraints before it takes effect. The authors evaluate the learned planner's behaviour in open-loop studies on real-world urban data, discuss the integration work required for stable closed-loop operation, and report results from deployment on their research vehicle.
3 more specialized papers
- Forward Trajectory Steering for Hamilton-Jacobi Reachability Analysis Sungje Park, Stephen Tu
- HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting Xikai Sun, Cangtian Zhou, Kebin Liu et al.
- DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation Yan Deng, Fei Xu
Reasoning 9
Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing
Reliability methods for large language model (LLM) reasoning are typically evaluated only by their own authors, so this work independently reproduces and stress-tests two of them: RPC, which aggregates token probabilities with self-consistency at inference time, and LCF, which trains projectors to split hidden states into content and logic parts and edits the logic part toward a valid region. Extending both methods to text-to-SQL, legal extraction, fallacy identification, and precedent grading, the authors find that RPC reproduces the original results exactly on the authors' released data but shows no statistically significant advantage over plain self-consistency on any new domain, with its largest apparent gain reversing sign when the evaluation sample is enlarged. LCF's logic-validity direction exists but is weak compared to a semantic control, and its intervention either has no significant effect or significantly hurts performance on most models tested.
Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets
A dense pretrained language model (Qwen2.5-0.5B-Instruct) is retrofitted with recurrent depth by splitting it into a prelude, a weight-tied recurrent block, and a coda, so that it learns an iterative latent transition computing one task step per loop. The mechanism installs at both a 6M-parameter adapter budget and a 180M full-block budget, persists when only final answers are graded, and extrapolates to roughly 1.5 times its supervised depth while holding 70% accuracy through depth 18. Against a same-size scratchpad-trained model, the recurrent variant wins overall at 84% versus 72%, retains 53% versus 2.5% accuracy beyond depth 10, and answers 7.6 times faster. A second task running the rule in reverse exposes a limit: the inverse was learnable in isolation, but no continuation acquired it while preserving the installed mechanism and general capability, marking a catastrophic-interference boundary.
Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning
Self-consistency — sampling N chains of thought and returning the plurality answer — is a standard way to spend inference-time compute, but on the 198-question GPQA Diamond benchmark of graduate-level science it reduces per-problem accuracy on most problems for two small instruction-tuned models: 56.6% of problems for Qwen2.5-7B and 65.7% for Llama-3-8B, findings pre-registered on a confirmatory split with all four hypotheses passing. An oracle that routes each problem to its best sample count sits 14-17 accuracy points above single-sample performance, but no verifier-free gate based on plurality agreement or token entropy captures any of that gap. The mechanism is that confidence does not track correctness: in the highest-agreement bin, Qwen's plurality answer is right only about half the time, and Llama's highest-agreement bin is less accurate than its lowest. Reasoning-native models were not tested and are flagged as the central open question.
Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs
Getting large language models (LLMs) to reliably verify and fix their own mistakes remains a fundamental challenge. Self-Fix Step-DPO (SFS-DPO) is a reinforcement-learning-based two-stage framework that first strengthens step-level reasoning through step-level preference optimization and then explicitly trains models to self-verify and self-correct; a teacher-assisted variant, SFS-DPO-R, adds explanatory rationales for error verification to provide stronger corrective signals. In-domain and out-of-domain evaluations across multiple LLMs show consistent gains over prior step-level training baselines, with analysis attributing the improvement to more frequent and more effective self-correction.
Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
On-policy distillation (OPD) is commonly believed to let a student language model acquire new reasoning abilities from a stronger teacher. The authors test this through test-time scaling, varying the sampling budget K and comparing pass@K and avg@K between OPD-trained models and their pre-OPD base models. OPD models keep a consistent avg@K advantage, but the pass@K edge shifts to the base models as K grows, and a problem-level analysis using pass@1024 shows OPD makes more previously solvable problems unsolvable than the reverse. The authors conclude that OPD behaves like an illusory distillation, improving sampling efficiency rather than genuinely expanding the student's reasoning capability boundary.
Policy-as-logic for robust reasoning over rules
Generative AI systems answering questions about tax rules, airline baggage allowances, and similar domains must respect written policies, but purely prompt-based approaches are brittle. The proposed hybrid symbolic method expresses policies in formal logic, uses the language model at inference time only to extract facts that ground the logic's predicates, and delegates the actual reasoning to an answer set solver, making responses interpretable and auditable. This separation of extraction from reasoning outperforms policy-as-prompt and policy-as-code methods in most cases with roughly a tenfold reduction in token usage, and stays accurate and robust under input perturbations.
OEIS Open: How many conjectures can language models turn into theorems?
The benchmark collects 492 open mathematical conjectures from the On-Line Encyclopedia of Integer Sequences (OEIS), formalized in Lean, with open-source evaluation code that runs any generic language model against them and is hardened against cheating attempts. Language models equipped with a minimal set of tools resolved 147 conjectures (30%) at a budget of $50 per attempt, and the best current model scores 44% on a 100-conjecture Lite subset at $200 per attempt. Giving models access to 476,000 arXiv papers or more sophisticated agent loops did not improve performance, and while the conjectures are of uncertain mathematical significance, the results show language models can resolve open research conjectures autonomously at modest cost.
2 more specialized papers
- Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library Xining Xun
- Latent variable models for simultaneous EOV identification and removal in population-based SHM M. D. Champneys, M. R. Jones, A. J. Hughes et al.
Unclassified 2
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization
No summary available — see the abstract on arXiv.
1 more specialized paper
- A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases Danial Sharifrazi, Saadat Behzadi, Nouman Javed et al.