Wednesday, September 2, 2026
Highlights
UI-Venus-2 Technical Report
Graphical user interface agents tend to be tuned for benchmarks rather than deployment, held back by narrow environment coverage, brittle task construction, and reward signals that cannot be trusted. UI-Venus-2 scales three axes together for a single closed-loop reasoning-and-action foundation agent spanning mobile, web, and desktop: coverage of more than 170 multilingual mobile apps plus native desktop operating systems, a deep-research pipeline that generates function-grounded instructions, and trace-level plus sample-level verifiers using visual keypoints and multi-model voting to produce reliable reinforcement learning rewards. Safety-aware mechanisms gate consequential actions, and the model is released open source.
Multimodal GUI agents score well on benchmarks but stay brittle in real deployments because environment coverage is narrow, generated tasks are poorly grounded, and reward verifiers are too coarse for reliable reinforcement learning. Ant Group's UI-Venus-2 (9B and 27B, initialized from Qwen3.5-9B and Qwen3.6-27B) attacks all three at once by jointly scaling environments, function-grounded task generation, and multi-model verification, and by unifying mobile, web, and desktop control in a single open-source policy.
- Training follows a three-stage recipe: large-scale trajectory mid-training, step-level offline RL run separately per domain (
Grounding,CAPTCHA,Mobile,Web,Computer), and a multi-teacher on-policy distillation (MOPD) stage that merges the domain experts while conditioning the distillation signal on action correctness, suppressing supervision when the action is right, emphasizing parameter tokens when only they are wrong, and masking parameters when the action type is wrong, with teachers additionally given a hint about the correct action type that the student never sees. - Data comes from a closed-loop pipeline that builds per-app capability catalogs via deep research (170+ mobile apps, 4,000+ web domains across 19 categories, desktop tasks serialized as
TaskSpecsnapshots), and verification uses Semantic Guided Verification, which extracts visual keypoints from the task, judges them window by window with a VLM, and aggregates votes across heterogeneous models into completed/partial/infeasible/failed labels used for data stratification rather than as raw rewards. - Headline results for the 27B model: 93.4% on
WebVoyager(595-task refresh), 80.2% onREAL, 84.0% onAndroidWorld, 80.5% onOSWorld-Verified, 55.5% onDeskCraft, and 79.9% Pass@1 on the newVenusBench-CAPTCHA(versus 53.0% forQwen3.6-27B), with the 9B model typically within a few points. - Grounding is strong but not uniformly state of the art: 80.1% on
VenusBench-GDis a new best, whileScreenSpot-Pro(74.1%) andUI-Vision(66.9%) trailQwen-UI-Agent-27B, which also leads onMobileWorldat 50 steps (82.1% vs 76.1%). - Long-horizon desktop work remains largely unsolved: on
OSWorld 2.0the 27B model reaches only 2.8% binary accuracy and 13.2% partial score against 13.0% forGPT-5.5, and the authors caution that many baseline comparisons use model-specific action scaffolds, different task subsets, and live websites whose state varies by evaluation date.
Safin-1: Safety from Within through Memory-Native State Evolution
Safety in foundation models is usually imposed through external safeguards or post-hoc alignment such as supervised fine-tuning rather than living inside the model's own computation. Safin-1 pursues the alternative through MARCH (Memory-Anchor Routing across Context History), an architecture that maintains structured memory states and retrieves relevant history via content-conditioned routing, supporting test-time adaptation of persistent capability states — including a dedicated Safety State — without repeatedly modifying the backbone. The authors report substantial safety improvements from state-based adaptation alongside evaluations of general capability, long-context understanding, retrieval, and efficiency, and describe the work as an initial architectural exploration rather than a completed programme.
Recurrent and linear-attention models compress the whole prefix into one evolving state, so earlier associations get overwritten and become unrecoverable, and safety behavior is usually bolted on through fine-tuning or external guards. Safin-1 addresses both with MARCH (Memory-Anchor Routing across Context History), which checkpoints the recurrent state into addressable anchors that tokens can route over, and reuses that same routed state bank to host a learned, detachable Safety State so safety is invoked through the model's native computation rather than imposed from outside.
- Every 512 tokens the recurrent state is snapshotted as a matrix-valued anchor paired with a content-derived 64-dim routing key; each token softmax-routes over causally visible anchors plus a learned zero-payload null option, and the weighted readout is added to the current-state path without altering the base recurrence, with a fused FlashAttention-style reader and
Top-4sparse routing that more than doubles training throughput versus dense routing at 128K tokens. - In matched 0.8B pretraining on 50B tokens,
MARCHonGated DeltaNetlifts the 8-task commonsense average from 40.1 to 41.5 (slightly above both full-attention baselines),LongBenchfrom 11.9 to 14.9 (+25%), real-world in-context retrieval from 20.5 to 23.3, and beats the best recurrent baseline in 19 of 24RULERNIAH settings, with the largest gains on multi-needle tasks and at 32K beyond the 16K training length, and the improvements hold acrossGDN,KDA, andGDN2backbones. - Scaling via matched continual pretraining and SFT from
Qwen3.5-4BandQwen3.5-35B-A3Braises the ten-benchmark macro average from 66.79 to 69.20 and 76.25 to 78.35, with gains concentrated on hard reasoning (AIME 2025Avg@64 up 8.0 and 8.2 points,MMLU-Pro+9.3 at 4B,GPQA-Diamond+6.6 at 35B-A3B) whileGSM8K,MATH, andMMLUstay within about a point. - With the backbone frozen, a persistent per-layer
Safety Statetrained on 1,915STAR-1/STAR-benignconversations (each replicated with 0, 512, and 2,048 tokens of benign prefix so it works across anchor-bank configurations) cuts average jailbreak attack success rate by 42.3% at 4B and 52.3% at 35B-A3B acrossWildJailbreak,FORTRESS,StrongREJECT,Jailbreak-R1, andJailbreakBench, with substantially lessXSTestover-refusal than a training-matched rank-8 LoRA. - The safety-state results table is truncated in the available text, so absolute ASR, over-refusal, and capability-retention numbers cannot be checked, the baseline ASRs are already low (roughly 6-7%) so the relative reductions rest on small denominators, dense routing beats
Top-4on NIAH in the 0.8B ablation even though the scaled models useTop-4, and the authors themselves frame this as only an initial architectural exploration of "Safety from Within."
EM^2Mem: Event-Centric Multimodal Memory for Large Language Models
Memory systems for long-video question answering usually store captions, frames, transcripts, and graph facts as separate fragments, forcing the language model to reassemble cross-modal and temporal alignments at inference time when context is scarce. EM^2Mem instead binds heterogeneous evidence to event anchors while memory is being built, so each event-indexed cell already carries aligned multimodal records, temporal context, graph relations, semantic facts, and provenance. Across three long-video QA benchmarks it gains 2.0, 2.4, and 3.7 accuracy points over the strongest memory baseline and 7.0 points of strict event-level top-5 evidence recall, while cutting per-query latency 4.67 times and total inference tokens by 63.66%.
Long-video question answering needs an external memory, but existing multimodal memories store captions, frames, transcripts, and graph facts as separate fragments that the LLM must re-align across modalities and time at inference, when context is scarce and attribution is hardest. EM^2Mem instead binds all heterogeneous evidence to short event anchors during memory construction, so retrieval returns grounded, generation-ready multimodal events rather than pieces to be stitched together.
- The video is cut into 30-second segments, each anchored to an event cell holding a keyframe caption, transcript, representative keyframes, and structured fields for actions, objects, topics, scenes, and entities, with multi-scale context views at 3-minute, 10-minute, and 1-hour spans, plus an episodic graph of cross-event entity and temporal links and a semantic graph of recurring habits and relations.
- At query time a lightweight retriever picks candidate event cells, expands them through graph links, an LLM selector filters them, and the answer model reads a compact evidence view with up to three keyframes for visual verification.
- Against the reproduced
WorldMMbaseline under the same setting, average accuracy rises 66.0 vs 64.0 onEgoLifeQAand 76.8 vs 73.1 onVideo-MME (L), and 67.7 vs 65.3 onEgo-R1 Benchagainst published numbers, while per-query latency drops from 459 s to 98 s (4.67x) and total inference tokens fall 63.66%. - Strict 30-second event-level Top-5 evidence recall reaches 30.8%, +7.0 points over
WorldMMafter five retrieval rounds, and ablations show temporal context views matter most (removing them costs 5.6 points), with construction-time unification beating retrieval-time fusion by 3.2 points on structured fields. - Margins over the originally published
WorldMMnumbers are thin (0.4 and 0.2 points onEgoLifeQAandVideo-MME (L)),WorldMMstays stronger on several habit and temporal categories, converting visuals to text fields loses fine pixel detail, upstream captioning errors propagate into memory, and the heavy upfront construction cost only breaks even on wall-clock time after roughly 23–24 queries per video.
Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
In multi-agent LLM pipelines a single prompt usually carries two entangled jobs — producing task content and specifying execution protocol such as message routing, output format, and termination signals — so an optimizer tuning the content can silently break the protocol and crash the pipeline. The proposed control-data flow separation encodes execution-critical control as typed, validated program objects while leaving only natural-language task content exposed to prompt optimization. Across synthetic reasoning, collaborative review generation, and insurance rating workflows, the framework achieves 100% eventual protocol validity while still improving task performance.
Prompt optimizers such as TextGrad and DSPy treat every line of an agent's prompt as editable, but in multi-agent pipelines those same prompts also encode the routing commands, output formats, and stop signals the surrounding Python controller parses. The paper proposes control-data flow separation: each agent emits a typed, schema-validated control object that only the controller reads, plus a free-form data message that other agents and the optimizer read, so prompt edits can never reach the execution interface.
- Control schemas are declared as Python dataclasses or Pydantic models with
Literalfields for closed sets such as routing targets, the auto-generated JSON scaffolding lives in a frozen prompt slot the optimizer cannot touch, and parse or validation failures trigger bounded retries or a default action rather than ever reaching the router, which the authors formalize as a protocol-stability lemma and ship as thecdseplibrary. - Across four settings (a
BBHsubset,MARGreview generation, and synthetic plus industry-verified insurance underwriting) the method reaches 100% episode stability and the best task score everywhere, including 78.3% BBH accuracy versus 74.3% forDSPy + BootstrapFewShot, 44.4 Jaccard on MARG versus 43.2 forDSPy + MIPROv2, and 36.7% on partner-rated underwriting versus 31.7% for the partner's own hand-written prompt. - Naive
TextGradcollapses on the routing-heavy tasks, falling to 0% stability on MARG and 56.7% on industry underwriting because the optimizer rewrites inline JSON and chapter-name instructions, and a prompt diff shows it touches control-relevant tokens about four times more often than the separated variant (16.6% versus 4.2% of edited lines on review). - Ablations attribute the properties to distinct components: schema scaffolding alone lifts MARG stability from 0% to 100%, parse retry adds only about one point of remaining reliability, and per-example feedback (rather than a scalar batch loss) drives the quality gains, raising review Jaccard from 26.9 to 38.0 and underwriting accuracy from 37.8% to 51.1% at the same budget.
- The guarantee covers only protocol validity, not semantic correctness, MARG scores rely on a
gpt-5.4-miniLLM judge rather than the original GPT-4 setup so are not comparable to the published benchmark, the cross-family runs on Claude and Gemini use a 12-paper subset with a single seed, and the framework assumes fixed agent roles and schemas with no support for runtime agent creation or schema evolution.
DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
Self-play lets a model generate its own training questions, but without direction its solver plateaus: existing unguided signals such as difficulty, learnability, or diversity keep questions hard and varied without saying which reasoning weaknesses to attack, while guided methods import direction from human examples or document corpora outside the loop. DiagEvo derives that direction from the solver's own failure history — a diagnostician extracts recurring error causes into a hierarchical memory that groups them under skill nodes and marks each Active or Mastered by self-consistency on targeted questions, and the challenger uses those states and recurrence counts to trade off cause-targeted generation against free exploration, with double-confidence filtering keeping mid-difficulty questions only when the majority answer leads clearly. With a 4B diagnostician it beats every baseline in mean accuracy across all nine benchmarks for Qwen3-4B, Qwen3-8B, and OctoThinker-8B, reaching 72.3% mean accuracy on five mathematical reasoning benchmarks with Qwen3-8B, 4.5 points above R-Zero.
Self-play curricula for reasoning models tend to stall because difficulty and diversity signals say how hard a question is but not why the solver fails. DiagEvo mines the solver's own failed trajectories for recurring error causes, stores them in a hierarchical memory with Active and Mastered states, and lets that memory steer the challenger toward the weaknesses that still need practice, without any external corpus, human examples, or difficulty labels.
- A frozen
Qwen3-4B-Instruct-2507diagnostician compares a failed GRPO trajectory against a pseudo-label-agreeing one to extract the earliest reasoning difference as a transferable cause, deduplicates it via embedding retrieval, and files it under a skill node; a cause is promoted to Mastered when mean self-consistency on questions targeting it reaches 0.70, and reactivated when new failures match it. - The challenger mixes cause-targeted generation (sampling Active causes in proportion to their failure counts, optionally stitched with a Mastered sibling from the same skill node) with free exploration, with the mixing odds scaling linearly with the normalized failure count so that mastered causes push probability back toward exploration.
- Double-confidence filtering keeps only questions whose majority vote share lies in 0.25 to 0.75 and whose top answer leads the runner-up by a ratio of at least 1.6, which raises oracle agreement of retained pseudo-labels from 65% to 75% at round 5 versus unguided self-play.
- On
Qwen3-8B-Baseit reaches 72.3% mean over five math benchmarks and 57.4% over all nine, 4.5 points aboveR-Zeroand 1.1 aboveDARC, while also topping every baseline onQwen3-4BandOctoThinker-8B; scaling the diagnostician to 235B adds only about 1 point, and the added diagnosis machinery costs 5.7% of wall-clock time. - Gains still plateau and reverse after round 5 (72.3% falls to 71.4% by round 7) with the round count fixed in advance, the filter measures agreement rather than correctness so shared solver errors can pass, and the curriculum is math-only with general-reasoning gains reported as transfer rather than direct construction.
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
Looped transformers add effective depth by iterating a shared block, but comparing at fixed model size hands the looped variant extra floating-point operations, conflating architecture with compute. SMELT loops the middle half of layers twice in a Mixture-of-Experts transformer while matching per-token FLOPs, total non-embedding parameters, and key-value cache against an unlooped baseline, scaled across four sizes up to 54B non-embedding parameters with a separate Chinchilla-style scaling law fit per architecture. Loss falls faster with compute, saving 6.8 to 18.0 percent of training FLOPs on the compute-optimal frontier, with downstream gains exceeding what validation loss predicts, largest on code, and growing with sequence length and in-context example count. Mechanistic analysis credits the second visit with shrinking the attention sink and redirecting attention toward content-relevant tokens.
Looped Transformers add effective depth by re-running a shared block of layers, but prior comparisons held parameter count fixed and let the looped model spend extra FLOPs and KV cache, so the architectural benefit was never isolated. The authors use Mixture-of-Experts to match per-token FLOPs, total non-embedding parameters, and KV cache simultaneously, then ask whether looping still wins across a full scaling ladder.
- The
SMELTrecipe loops the middle 50% of layers twice, pays for the second pass by narrowing the hidden dimension, restores capacity by raising expert count (192 to 288 at the 200M scale), scales looped residual updates by 1/2, and shrinks head size at a higher GQA ratio so all three budgets land within about 4%; ablations at 200M showed a 50% span beats full-stack looping, two visits beat three or four (which force a thinner model), and the looped model tolerates a larger effective depth-to-width ratio than the unloopedBaseline. - Across a 4×4 grid of scales (100M to 1.6B active, up to 54B non-embedding parameters) and compute-equivalent sparsities of 85/95/97%, separately fitted Chinchilla-style surfaces give
SMELTa steeper frontier exponent (γ = 0.250 vs 0.237), translating to 6.8–10.0% training-FLOP savings at 10²⁰ FLOPs and 14.7–18.0% at 10²¹, while compute-optimal tokens-per-parameter stay within 6% of theBaseline, so the saving comes from lower loss at the same allocation rather than reallocating budget. - Downstream,
SMELTwins 96/96 matched pairs onDCLM Completion, 83/96 onDCLM Core, and 29/30 above-chance pairs onMMLU, and its residual above a sigmoid calibration from validation loss to benchmark score is positive at every scale and grows with model size, with Code the most-improved training domain and gains that widen with sample length and number of in-context examples. - Mechanistic probes show the second visit selects largely the same experts and attends largely the same tokens as the first but produces larger, aligned residual updates that amplify rather than overwrite, changes values more than queries and keys, and reduces attention-sink mass in favor of content-relevant tokens.
- Caveats: every configuration was trained once so reported error bars exclude training-run variance, the cell-bootstrap intervals are wide (10²⁰ FLOPs at 85% sparsity spans [1, 22]%), the 10²² FLOP gains of 19.6–23.5% are extrapolated beyond the fitted window with intervals reaching zero, the dense S=0 control was excluded from the fit, and the data and
Baselinefamily are internal so the loop-span choice rested on validation loss becauseDCLMmetrics did not track it in the ablation.
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
How far multimodal large language models (MLLMs) extend from perceiving to acting is tested by dropping one directly into a drone's control loop with its entire action space declared solely in the prompt. DroneCATS-Agent makes the model a swappable component and the DroneCATS benchmark treats it as the independent variable across approaching a visible target, tracking a moving one, searching outside the initial view, and commanding a multi-drone fleet, with no fine-tuning or function-calling schemas and a roster scaling down to 2B parameters. Small open models often navigate into the success radius more reliably than frontier models yet lose the episode by declaring arrival prematurely or never at all, and multi-drone commanding widens the gap as small models blindly copy one coordinate across distinct views, making protocol discipline and correct termination, not navigation, the bottleneck.
Prior MLLM drone-control systems (TypeFly, SPF, Fly0, OnFly) progressively strip decisions away from the model, which is why backbone choice appears not to matter; this work holds a minimal agent fixed and makes the model the independent variable. DroneCATS-Agent declares four actions (go, rotate, think, finished) purely in the prompt with no fine-tuning or function-calling schema, and the DroneCATS benchmark scores approaching, searching, tracking, search-and-track, and four-drone commanding under one criterion anchored on the model's own arrival declaration.
- Each step the model sees only an egocentric monocular RGB frame, the instruction, and its last five actions, and returns one JSON action:
gogives a pixel plus a self-estimated depth that a rule-based controller lifts to a velocity setpoint,rotateyaws in place,thinkhovers and unlocks reasoning for the next call, andfinishedis a claim the scaffold logs while the episode keeps flying, judged post hoc as success only if some declaration occurred within 5 m of the target while it was visible (100 AirSim episodes, 300 s cap). - Even the simplest cell is unsolved: the best model,
Gemini 3.7 Flash, hits 65% on approaching, 40% on searching, 80% on tracking, and 45% on search-and-track; no model exceeds 40% when the target is withheld from the first frame;GPT-5approaches at 60% but tracks at only 15%; the embodiment-specialisedGemini Robotics-ER 2averages 47.5% versus its generalist sibling's 57.5%; and theQwen3.5ladder falls monotonically 33.8/21.3/12.5/0% from 27B to 2B. - The headline finding is that small models fail on protocol, not navigation:
Qwen3.5-9Benters the success radius in 90% of approaching episodes, more than any frontier model, yet converts only 35% by declaring at 0.63 of its start distance, whileQwen3.5-2Bdeclares at 1.28× its start distance and never succeeds, andCosmos3-Edge-2Breaches the radius in 25% of episodes but never declares at all. - Commanding four drones from one context toward look-alike targets distinguishable only by a licence plate or sign amplifies this split:
Gemini 3.7 Flashwins 80% of episodes andGemini Robotics-ER 265%, butGPT-5drops to 20%,Claude Opus 5to 0%, andQwen3.5-9BandQwen3.5-27Bpaste an identical point across all four distinct views on 70% and 58% of go-steps, so at most one command can be grounded. - Caveats are stated plainly: 20 episodes per cell yields a per-cell standard deviation of 9–13 points (a three-flight audit gave 54.6 ± 9.7% overall), so orderings within a tier are within noise; results are simulation-only; success degrades when the control loop slows below ~2 Hz; the
thinkaction was invoked 1,797 times but measurably changed nothing (+0.03 m/step, reacquisition 13.5% vs 13.4%); and hierarchical commander-plus-subagent control is left unevaluated.
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
Autonomous software development asks LLM coding agents to turn high-level requirements into complete working systems without human intervention. Harness-of-Harness (HoH) wraps existing coding-agent harnesses in iterative planning-coding-testing loops, balancing repair against capability growth, scoping work into small verifiable increments, separating implementation-time testing from independent evaluation, progressively exposing deliverables, role-specific tools, and skills, and maintaining versioned project history. Across GameCraft-Bench, FrontierSWE, and ProgramBench with three harness-model pairs (Codex with GPT-5.5, OpenCode with DeepSeek-V4-Pro, and Pi with MiniMax-M3), it beats the standalone harnesses by an average relative gain of 52.25 percent after three iterations, peaking at 82.86 percent. A multi-day deployment of more than 70 iterations produced a playable first-person-shooter game with a coherent storyline, implemented core mechanics, visuals, and audio.
Coding agents work well within a bounded episode but struggle to carry a from-scratch software project forward over many sessions, losing track of requirements, unresolved failures, and already-verified behavior. Harness-of-Harness (HoH) wraps an unmodified coding harness in a repeated planner–developer–tester loop that hands both the evolving artifact and structured test evidence from one iteration to the next.
- Each loop invokes the same harness–model pair three times in fixed roles: a
Project Plannerreads the spec plus prior test evidence and writes a bounded development document, aDeveloper(the sole writer) implements it with shift-left self-tests, and a read-onlyQA Testerevaluates a frozen candidate with black-box and white-box checks and returns a structured report; outputs that violate the required schema trigger a retry, while reasoning and tool use are left unconstrained. - After three iterations,
HoHbeats the standalone harness on all three benchmarks for all three configurations (Codex+GPT-5.5,OpenCode+DeepSeek-V4-Pro,Pi+MiniMax-M3), with an average relative gain of 52.25% and absolute gains of 16.62–22.08 points onGameCraft-Bench, 19–29 points of dominance onFrontierSWE, and 6.09–16.85 points of test pass rate onProgramBench; extendingCodexto ten loops onFrontierSWElifts dominance from 39.33% at HoH@3 to 72.67%. - The gain is not just extra compute: at matched pass counts on
GameCraft-Bench,HoHoutscores repeated-sessionVanilla Continuationby 10–13 points, and HoH@2 at 5.67M tokens (64.84) beats three-pass continuation at 6.33M tokens (58.24); ablations show freezing the plan, dropping test evidence, or rebuilding from scratch each loop costs 6.28–8.13 points. - In a multi-day run with
CodexandGPT-5.6-Sol, plus Godot MCP tooling and asset, UI, and testing skills,HoHbuilt a playable narrative first-person shooter over 70 loops, closing 65 of 81 recorded issues, though 16 remained open and 17 issues were reopened after later changes regressed verified behavior. - Benchmarks were sampled (45 of 140
GameCraft-Benchtasks, 15 of 17FrontierSWEtasks) and the multi-day case is a single project with one configuration, so generalization beyond game development and the tested harnesses rests on limited evidence, andPipeaked at HoH@2 rather than HoH@3 onProgramBench.
H3-World: Turning Language Understanding into World Control
H3-World turns the 33B MiniMax-H3 video generator into an interactive world model by exploiting the observation that large video generators already accept coarse natural-language control of character behavior and camera motion. Each action is represented as a structured combination of character and camera instructions aligned to the corresponding temporal video latents, and temporal attention routing confines each instruction to its intended time interval to reduce control leakage across actions, with no dedicated action modules added. Using only 8,000 gameplay samples, 10,000 LoRA optimization steps, and 0.199% trainable parameters, the system achieves effective character and camera control, preserves generation quality, and generalizes to unseen scenarios.
Interactive world models normally need action-conditioned video generators built with dedicated control modules trained on aligned action-video trajectories. H3-World instead observes that the 33B MiniMax-H3 video generator already follows coarse language instructions for character and camera motion, and turns that into precise, temporally grounded control by expressing each action as text bound to a specific video latent interval and adapting only 0.199% of the parameters with LoRA.
- Each discrete control state (eight character and camera keys plus a camera-speed flag) is mapped to a short compositional instruction such as "the man walks backward and strafes left, camera pans right slowly," drawn from 9 character clauses and 16 camera clauses, and injected through the model's native text pathway rather than a learned action embedding.
- Every video latent interval gets its own independently encoded action prompt with a mirrored temporal position, and a deterministic single-egress attention mask lets each action span be read only by its matched latent, so actions enter the video stream at one point and then propagate through the unchanged bidirectional video-to-video attention.
- Training uses only 7,872 gameplay clips from
ABot-World-Explorer-500h(124 frames, 832×480, 37 action prompts per clip) with rank-32 LoRA for 10,000 steps on the attention projections and a two-layer token refiner, keeping the backbone, encoder, and VAE frozen. - On a controlled left-then-right pan schedule,
H3-Worldproduces cumulative horizontal flow of +52.7 before and -106.0 after the switch, while frozen H3 with a global prompt gives 0.0 and -17.3 and a zero-LoRA per-latent interface stays nearly static, showing that both the temporal binding and the adaptation are needed; text-based control also beats additive-bias and FiLM action conditioning in qualitative comparisons. - The model composes unseen character-camera pairs (52 of 135 valid combinations never appear in training) and transfers to out-of-distribution scenes, but evaluation is mostly qualitative on representative examples, generation is fixed-length and short-horizon, and there is no persistent world state, real-time interaction, or planning.
StudentSim: Training LLM-based Student Simulators
AI tutors work best when they adapt to individual students, but evidence about which guidance suits which learner is slow and costly to gather, and existing student simulators either track state without processing explanations or role-play fluently without matching the target student's competence. StudentSim turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization, so each simulator both mirrors a student's own responses and updates them under tutor guidance. The accompanying StudentSimEval protocol covers 60 students across chess, second-language English writing, and mathematics and measures behavioral fidelity and guidance responsiveness; StudentSim outperforms GPT-5.4 on both metrics in all three domains, reaching fidelity 0.51 and responsiveness 0.91 in chess versus 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. Used as a reward model for tutor reinforcement learning, it produced a chess tutor that expert humans rated more accurate, better-guided, and more personalized than both a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward.
Real-student feedback on which tutoring moves work for which learner is sparse and slow to collect, and existing simulators cover only half the job: behavior-tracking models mimic a student but cannot read tutor guidance, while prompted LLMs follow guidance but fail to reproduce a specific student's competence. StudentSim trains one small simulator per student by first pooling records across all students in a domain and then specializing that base on each individual's sparse data, and StudentSimEval scores any simulator on both behavioral fidelity and guidance responsiveness.
- Both stages fine-tune LoRA adapters on
Qwen3-4B-Instruct, mixing single-turn records that teach a student's own responses with multi-turn records that pair a wrong answer, tutor guidance, and the canonical correction, so the model learns to update toward the fix rather than just imitate. - The benchmark covers 60 students across chess (Lichess games), second-language English writing (
EFCAMDAT), and multiple-choice math, and evaluates every method on the same frozen per-student held-out records. - In chess,
StudentSimreaches fidelity 0.51 and responsiveness 0.91, versus 0.23 and 0.72 for promptedGPT-5.4and 0.45 and 0.27 forMaia2, with the same ranking holding in L2 writing (0.56 / 0.64) and math (0.64 / 0.92); the gap is largest on Socratic and conceptual guidance, where the tutor never states the answer. - As a proof of concept, using the pooled simulator as a GRPO reward for a
Qwen3-VL-8Bchess tutor yields guidance that expert players rate at 90.5% factual accuracy versus 75.7% for the no-RL baseline and 71.6% for a tutor trained against aGPT-5.4simulator reward, with higher guidance and personalization scores as well. - Limitations include very small per-student evaluation sets outside chess (tens of records per learner), LLM-generated rather than real tutor guidance for chess and math, a tutor-RL demonstration confined to chess because it needs an engine-based reward, and the metrics capturing only a single-step update rather than learning, retention, or forgetting over time.
Large Language Models 105
InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information
Competitive-programming benchmarks for large language models (LLMs) hand over the full problem input up front, which skips the ability to reason when facts are revealed only in response to queries. InteractBench gathers 322 interactive problems from Codeforces, AtCoder, IOI, and ICPC, each packaged with an executable local interactor so generated programs can be judged fully offline across multi-round exchanges bound by protocol rules and query budgets. Even the strongest reasoning models achieve only limited success on these problems, and a fine-grained failure taxonomy traces the gap mainly to algorithmic logic errors but also to frequent protocol violations and exhausted query budgets.
Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning
Persona prompting usually relies on synthetic personas that flatten real variation and lean on stereotypes rather than the signals that actually drive preferences. Profile behavioral grounding instead derives open-ended, high-fidelity user profiles from authentic anonymized social media posts, then tests them both as training data for supervised fine-tuning and as non-parametric context for test-time multi-perspective reasoning. Behavior-derived profiles beat synthetic-persona baselines and improve base models on recommendation and open-ended query benchmarks under both paradigms, with code released publicly.
ES-AHD: An Evolution Strategy Framework for Automatic Heuristic Design
LLM-driven automatic heuristic design typically mutates individual programs at random, which searches blindly and balances exploration against exploitation poorly. ES-AHD borrows two ideas from evolution strategies: semantic recombination replaces point-to-point reproduction by having the LLM extract shared insights from top-performing individuals to define a search direction, and the covariance matrix is mapped onto the model's sampling temperature so the search radius shrinks for local code refinement while occasionally spiking to escape semantic local optima. The result is a directional, center-guided sampling scheme that the authors report accelerates the discovery of high-quality heuristics, with source code released.
REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent
Post-training quantization methods solve each layer with one closed-form second-order step, which forces heavy approximation of the global loss — dropping cross-channel coupling, pooling output rows — and then freezes the resulting Hessian for the whole layer even as the loss landscape shifts column by column. REAL-Q keeps an end-to-end-aligned surrogate of the global loss and refines it with block-wise gradient descent after every 128-column block, adding a sliding window across layer boundaries to limit error propagation. On LLaMA-3.1 at 8B and 70B and Qwen3 from 0.6B to 32B at 4-bit weights with 16-bit activations, it cuts end-to-end KL divergence by up to roughly 49% against state-of-the-art globally guided methods.
Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning
Fine-tuning can erode in-context learning (ICL), and a common diagnostic assumes a model is still context-sensitive if its attention shifts when the demonstrations change. Formalizing that proxy as In-Context Sensitivity (ICS) — the average row distance between last-token attention on matched versus mismatched demonstration prefixes — and pairing it with the behavioural accuracy gap, a four-arm ablation on Llama-2-7B shows an ICS-maximising regulariser pushing ICS to 1.413, within 0.5% of its geometric ceiling, while the behavioural gap stays near zero and MMLU accuracy falls from 0.371 to 0.279. Endpoint analysis finds attention becoming sharp and near-disjoint across prefixes but routing to formatting and demonstration-body tokens rather than labels, a textbook Goodhart failure; objectives anchored to the pretrained computation instead hold a high-MMLU, moderate-ICS region.
OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization
NVFP4 is a microscaling format for low-bit inference, but a single large activation in a quantization block dominates the shared block scale and inflates the error of every other value in that block — an effect the authors name Collateral Quantization Error. OCGQuant attacks it purely through channel grouping, adaptively pairing outlier channels with low-magnitude companion channels so blocks are composed more favourably, without the mixed precision, rotations, or residual compensation that other post-training quantization methods add. On Llama3 and Qwen3 it achieves the lowest WikiText-2 perplexity and the highest average downstream accuracy among the evaluated methods while keeping prefill speedup close to round-to-nearest and matching its peak decoding memory.
KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training
Continued pre-training struggles to install knowledge from niche documents such as manuals or technical specifications, because those sources rarely repeat a fact, and the usual remedy of generating many paraphrases requires an expensive teacher model. KItCAT (Knowledge Injection via Corrupted Auto-regressive Training) augments ordinary next-token prediction by randomly replacing a subset of input tokens with other vocabulary tokens while leaving the next-token labels unchanged, producing many varied training inputs per document at negligible cost. The corrupted-input training consistently outperforms standard continued pre-training across multiple datasets and model families, reducing the need for paraphrase generation in decoder-only models.
Commit-first LLM judging inherits the judge's own errors
One defence against systems gaming an LLM judge is commit-first judging, where the judge solves the task itself, commits to an answer, and accepts a candidate only if the two match. An audit of default judge configurations in eight widely used evaluation frameworks finds none of the 24 configurations in scope implement it, while nine use a variant the literature measures as ineffective, all traceable to a single ancestor prompt through a copied typographical error. In a controlled experiment, plain best-of-N search with no access to correct answers got 90 of 96 candidates accepted on an interval-merging task under one documented configuration, with every accepted program failing a held-out suite; commit-first judging dropped acceptance to zero there, but on a second task the judge's own committed answer was wrong and the population converged on it, showing the defence relocates the gameable anchor to the judge rather than eliminating it.
Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding
Long-context decoding in large language models is bounded by memory bandwidth and the quadratic cost of attention, and existing sparse-attention schemes trade the memory overhead of metadata indices against the compute cost of adaptive block selection. Faster Flash Decoding (FFD) fuses the selector and the computer into a single kernel, replacing external metadata with content-aware scanning under low-bit quantization, and adds a top-delta rule that filters blocks to a distribution-adaptive sparsity level without global synchronization. The training-free, drop-in design reuses scanning results for the subsequent computation and reaches up to 11.6x kernel-level speedup and 2.37x end-to-end throughput at context lengths up to 256K, with accuracy preserved on RULER and LongBench.
Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs
Claims that multilingual language models route computation through language-specific states such as an English pivot rest on probes that infer a latent language from different signals, either hidden state geometry or what can be decoded from intermediate representations. Comparing these probes across model families, training regimes, domains, tasks, checkpoints, and up to 27 languages shows they systematically disagree: the Gaussian mixture model representation probe finds cross-lingual mixing earlier in the network, while decoding-based probes retain sharper, more English-biased language-specific signals. The disagreement tracks how multilingual a model is and how far training has progressed while staying comparatively stable across domains, which the authors read as evidence that current probes expose different aspects of multilingual processing rather than a single internal lingua franca.
Do General NLP Embeddings Capture Ontological Reasoning?
Whether general-purpose text embedding models encode logic-sensitive ontological structure is tested with AVA, a benchmark of 171,007 contrastive triplets built from 163 heterogeneous ontologies via hierarchy inversion, relation substitution, and disjointness injection, each pairing a statement with an equivalent paraphrase and a hard negative whose relational meaning contradicts it. Across more than 25 current embedding models, the best reaches 0.739 triplet accuracy but only 0.135 on the hard negatives. Fine-tuning improves discrimination substantially yet transfers poorly to downstream Semantic Web tasks such as taxonomy discovery and ontology alignment, and further analysis suggests much of the gain comes from recognizing the specific perturbation patterns rather than from robust ontological understanding.
Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs
Because language models are trained on static corpora, their knowledge goes stale, and existing tests of knowledge editing either get contaminated quickly or rely on counterfactual edits that clash with facts the model already holds firmly. The authors build ParallelEvents, a benchmark of fictional but plausible future worlds that generates coherent event trajectories, sidestepping contamination while staying internally consistent, and pair it with Synapse, a training framework that updates parameters using model-generated data through mid-training and instruction tuning. Synapse beats existing knowledge-insertion methods by 14.23%, suggesting simulated synthetic data can integrate new knowledge at scale without hand-curated corpora.
LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts
Sociodemographic prompting conditions a language model judge on an annotator's demographic profile in hopes of reproducing that group's judgments on subjective tasks. Comparing the predicted label distributions of 23 open-weight models against real annotator groups on three tasks, under no demographic information, single-attribute profiles, and intersectional profiles over gender, age, race, and education, the authors find that an unconditioned judge is not neutral — it best matches White, college-educated annotators — and that demographic conditioning is asymmetric, moving predictions toward majority groups and away from minority groups, most sharply on offensiveness where intersectional profiles amplify the effect. Comparing base against instruct models points to instruction tuning as a likely source of the asymmetry, suggesting the technique often fails the minority groups it is invoked to represent.
QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization
Weight-only post-training quantization cuts the cost of serving large language models but degrades badly below 2 bits, and the usual fix — unstructured sparsity — sacrifices the regularity that makes GPU kernels fast. QTEA quantizes weights to ternary values and compensates with salient weights as residual error correctors, confining those residuals to selected columns under semi-structured 1:4 sparsity, and adds column-wise rescale refinement plus an error-decay term that counters the order-dependent error accumulation in GPTQ-style column-by-column quantization. On Qwen3-14B it compresses all weights to an effective 1.7 bits per weight while improving average accuracy 16.7% over the strongest ternary baseline, with similar gains on Llama3-8B, and a lookup-table kernel delivers 7.2× faster per-token generation than FP16.
Authority Bias in Conversational Search Engines for Academic Paper Recommendation
Whether large language models recommending academic papers judge on content or on prestige signals had not been tested causally. Holding title and abstract fixed, the authors varied author prestige, venue, and citation metadata across original, flipped, and boosted conditions for eight models (five open-weight, three closed) in single-turn top-1 recommendation. Authority bias proved substantial and directional, varying widely by model and only partly reducible through prompt-level debiasing. They also document a say-do gap: debiasing instructions suppress explicit mentions of authority much faster than they suppress authority-driven choice flips, so audits based on what a model says systematically understate its behavior.
Hypotheses-Guided Self Distillation for Continual Personalization
Personalizing an assistant over long-term use is hard because users rarely state preferences outright; they leak through heterogeneous, latent, noisy signals, and current approaches either stuff raw histories into context or run expensive reward-based optimization. HypReflect instead infers explicit, uncertainty-aware preference hypotheses from those signals, reflectively revises them as evidence accumulates, and folds the resulting user model back into the assistant through hypotheses-guided self-distillation. It beats raw-history and incremental-update baselines across online personalization, multi-session interaction, and implicit behavioral signals, and generalizes to unseen users and new domains while staying stable across context budgets.
Latent Mechanisms of Language Control in Multilingual Language Models
Multilingual models sometimes switch languages mid-generation without reason, which suggests identifiable internal features that control output language. Three ways of locating such features in cross-layer transcoders are compared: selection by activation value (ValSel), by activation frequency (FreqSel), and by LLM-generated annotations of what each latent does (AnnSel), evaluated on two new code-switching benchmarks spanning seven languages with intervention experiments on Gemma-2-2B and Qwen3-4B. All three steer generation language effectively, with FreqSel strongest overall and AnnSel offering human-readable selection criteria. A knock-out analysis shows the three methods pick largely non-overlapping yet individually sufficient latent subsets, pointing to redundancy rather than one canonical language direction.
The Curse of Multilinguality in Lexical Normalization
Rewriting non-standard text like tmrw or gr8 into standard forms suffers from scarce labelled data, so practitioners often train one model across many languages at once. Holding a character-level model at fixed capacity and sweeping the number of jointly trained languages from one to twelve on a standard benchmark, per-language accuracy peaks when a language is paired with only one to four others and then declines by roughly forty percent as the remaining languages are added. A control that keeps total training data constant makes the drop arrive earlier and go deeper, implicating capacity competition rather than data volume, and typological distance from the other languages gives no dependable rule for how many co-training languages are best.
Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance
Conformance suites for quantized matrix-multiply kernels check whether two implementations agree within a numerical tolerance, and this work measures what such a check can actually catch. Injecting nine faults into a reference INT8 pipeline across 8,232 layer-fault-regime cells of Qwen3-1.7B, all five epilogue faults, including scale precision, double rounding, multiplication order, output truncation, and fused ordering, move the output by at most one bfloat16 spacing, so a one-spacing tolerance is blind to the entire class by construction and misses four of the five outright. Faults that break the accumulator's exactness preconditions or operand sharing are always caught, so the suite establishes those properties rather than interchangeability. Requantizing every weight scale to the nearest power of two makes CUTLASS and Triton agree bitwise at every linear layer (196/196 and 252/252, versus 8/196 and 10/252 with the original scales) and produces byte-identical token sequences at 1.7B, 8B, and 14B, with perplexity shifts under 1% and a previously reported +157% regression traced almost entirely to a probe that rewrote scales without requantizing weights.
Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning
Data engineering uses of language models, such as translating natural language into SQL and automating spreadsheet operations, need both better accuracy without finetuning and relief from the quadratic cost of attention over long contexts. The authors propose a drop-in neurosymbolic layer that slots into existing model backbones to strengthen logical reasoning and compress context, reporting an average accuracy increase of 85% across benchmarks including BIRD-CRITIC and LiveSQLBench with no task-specific finetuning or RLHF. Applied to long-context inference, the same symbolic prioritization and compression cuts effective token usage by more than half and brings effective time complexity down from quadratic to roughly linear on certain long-context tasks.
Removable and Irreducible: A Token-Cost Ledger for the Multilingual Tokenization Tax
Non-English text costs several times more tokens for the same content, and because attention is quadratic in sequence length that inflates compute as well — but how much of the penalty is actually removable is unclear. Framing tokenization as source coding with a Shannon-rate floor, the authors assemble a token-cost ledger that splits each language's cost at fixed parallel content into removable coding redundancy, residual coding slack, an intrinsic-content term, and an irreducible grapheme-to-phoneme term. On FLORES-200 across eight languages, a production tokenizer charges up to 8.9x more tokens for Indic scripts than for English, yet a script-matched code trained on only 1,012 sentences removes a median 64% of that excess, and a script-fair information floor shows intrinsic content differs by under 6% — the tax is representational rather than informational, and implies up to 79x attention cost. The authors scope this explicitly as compute-and-memory accounting rather than a model-quality claim, and release a one-command harness.
DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference
Mixture-of-Experts (MoE) models scale large language model inference cheaply in arithmetic but move a lot of weight data, which becomes the bottleneck on neural processing unit (NPU) systems; near-data processing (NDP) can absorb some of it, yet existing NPU-NDP MoE systems ignore hardware heterogeneity, dynamic expert-level concurrency, and temporal expert reuse across a batch. DynaNDE adds an analytical performance model covering heterogeneous hardware, data-movement cost, and communication-computation overlap, uses it to decide per layer which experts run on the NPU versus near memory while accounting for expert concurrency, and pairs that with a reuse-aware runtime that skips reloading experts already resident in NPU memory. Against the state-of-the-art NPU-NDP MoE serving framework it reports average speedups of 2.6x for prefill and 2.2x for decoding.
Late Transformer Layers Recode Syntax Canonically: Evidence from Greek Scrambling and Cross-Layer Generalisation
Probing work has established that syntactic information is decodable from early and middle transformer layers, but what happens to it in later layers is poorly understood. Cross-layer generalisation analysis on three Greek-tuned large language models, using minimal pairs of Modern Greek object-relative constructions that differ only in canonical subject-verb-object versus non-canonical verb-subject-object order while preserving meaning, finds that a probe trained on layers 20-31 transfers below chance to each early layer individually, classifying 99.3% of non-canonical sentences as canonical. Probe coefficients reverse sign around layer 22, which the authors interpret as a directional recoding toward the canonical form rather than simple information loss. The result characterises a representational format change beyond the known decline in syntactic decodability, and yields a directly testable prediction for human EEG and MEG decoding on the same stimuli.
HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference
Block quantization quantizes both weights and activations of large language models so inference runs on one low-precision datapath, but its design space of bit-width, block size, scaling scheme, and numeric format is largely unmapped. A design-space exploration shows larger blocks amortize dequantization and accumulation cost yet hurt accuracy, motivating HBQ (Hierarchical Block Quantization), which keeps large blocks for efficiency and adds a cheap significand-based second-level scaling that compensates block error better than power-of-two or integer scaling. The accurate variant reaches W4A16-level accuracy at W4A5 in less silicon area than NVFP4, and a 28nm ASIC applying HBQ to weights, activations, and the KV cache delivers 2.3x area and 4.6x energy efficiency over state-of-the-art weight-only quantization at matched accuracy, plus 1.5-3.0x speedup over prior block quantization.
Can LLMs Use Relational Transformer Embeddings?
Feeding frozen relational-encoder embeddings into a language model as soft tokens looks like an elegant division of labor, letting the encoder handle multi-table structure while the LLM handles reasoning without lossy text serialization of the database. The test injects embeddings from a frozen Relational Transformer into Qwen3.5-4B through a learned MLP projection plus LoRA, trained with supervised fine-tuning on chain-of-thought traces and then group-based reinforcement learning (GSPO), evaluated on 10 binary classification tasks over 6 RelBench databases under single-task, within-dataset, cross-dataset, and all-task supervision. The hybrid fails to consistently beat the standalone relational encoder, is frequently below random, and proves highly sensitive to serialization format and relational-token budget as well as unstable under RL; the authors publish the negative result and argue soft-token fusion needs stronger alignment objectives and schema-aware design.
Toppling the Hierarchy in Byte-level Language Modeling
Byte-level language models still fail at seemingly trivial character manipulation, and the leading designs are hierarchical, starting from bytes, downsampling to word-like units, then upsampling back to bytes for efficiency. Comparing hierarchical variants against pure byte-level models shows the hierarchy itself is what limits character-level understanding, with flat byte models consistently winning on character manipulation tasks. Ablating transformer layers into attention and feed-forward parts localizes the effect to byte-level attention as the primary mechanism, establishing an explicit trade-off between computational efficiency and fine-grained character understanding.
Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
A gap of under a percentage point on a leaderboard is routinely read as one model being better than another, but that ordering may be an artifact of which benchmark items happen to be included. Using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item response theory, the authors identify items with low residual differential item functioning across model families in one owner-disjoint fold, then use those frozen, difficulty-balanced weights to rescore models in the other fold, with length-matched random subtests as a control. Overall rankings barely move (Kendall's tau-b of .900 to .948), yet in four of five benchmarks 30.9-47.1% of cross-family pairs initially within one percentage point flip order, 16.9-28.6 points above the matched-random baseline, so sub-one-point leaderboard gaps need explicit evidence that the ordering survives benchmark recomposition.
Human-Anchored Factuality Evaluation with Strategic Annotation
Automated factuality judges scale evaluation but drift systematically from human ratings, so the setting studied is a hybrid one: run the judge over the whole dataset, buy human labels for a small subset, and combine them into a statistically valid estimate under a fixed annotation budget. The core observation is that judge-human disagreement in factuality is not just low confidence but has structure — incomplete evidence, temporal mismatch, unverifiable claims, rubric misalignment — so the authors build an annotation policy from failure-space analysis (FSA) signals that predict where the judge and humans will diverge. Against uniform and uncertainty-driven sampling, the FSA-guided policy raises effective sample size by 40.3% on an internal reference-based system (AutoFA) and 27.1% on RAGTruth, both settings where the judge underestimates true factual accuracy.
ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) reduces hallucination but dense retrieval handles multi-hop question answering poorly, and graph-based RAG, which does follow multi-step relations, pays for it with semantic drift and slow global graph traversal at query time. ISO-RAG projects the knowledge graph into a hyperbolic Poincaré ball offline to precompute per-node isoperimetric profiles, then uses those profiles to prune spurious edges at retrieval time so Personalized PageRank diffusion runs over a strictly local subgraph and converges quickly. Across multi-hop question answering benchmarks this yields average absolute gains of 10.0% in retrieval recall and 4.3% in downstream exact match while removing the latency cost of global traversal.
The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space
The interlingua hypothesis proposed here says large language models translate by reading a source sentence into a language-agnostic latent feature space and then generating the target sentence out of that same space, rather than learning pairwise source-to-target mappings. Three lines of evidence are offered: BLEU variation across language pairs is largely predicted by per-language competence with no pair-specific interaction terms; many internal components are causally involved in both monolingual tasks and translation; and fine-tuning on monolingual data alone recovers a large share of the translation gains obtained by fine-tuning on aligned parallel documents. The authors argue this reframes how translation ability can be measured and improved in general-purpose models.
Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation
Retrieval-augmented generation draws on corpora that can contain outdated, contradictory, or unreliable documents, and human reliability labels are too expensive to collect at corpus scale. TrustPropRAG builds a graph of document relations and solves an optimization problem that combines pairwise relations with a small set of human feedback labels, propagating trust scores multiple hops so that a limited number of costly reliability judgments extends across the whole corpus. The resulting scores drive both document selection and trust-aware answer generation, improving retrieval quality and exact-match accuracy over baselines while staying robust when feedback is sparse or noisy.
EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection
A model whose choices vary with the input rather than tracking its own output prior is often taken as evidence of task competence. Using a forced-choice signalling task abstracted from the board game Deception: Murder in Hong Kong, where the fit-maximising, posterior-maximising, and uniform-random reference strategies are all computable in closed form, the authors test seven language models under three scoring rules and find every one of the 21 model-by-rule cells reliably item-sensitive — yet 8 of those cells are statistically indistinguishable from random choice and 5 score worse than random, with item-sensitivity and distance from random correlating at only r = 0.30. They call this consistency without alignment and argue it undermines any evaluation resting on item-sensitivity, permutation consistency, or self-consistency without an independent reference; a literal-similarity baseline with no pragmatic reasoning beats most of the tested models.
Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs
Mixture-of-experts models are often shrunk by residual sparsification, which splits every expert projection into a shared base matrix plus a per-expert residual and then compresses each residual to minimize its own reconstruction error. The authors show that per-matrix error is the wrong objective, because an expert's output couples several projections and hidden representations, so small isolated errors compound into large output errors. PARSER retargets compression at expert output error using an output-importance measure of each weight's actual contribution, narrowing the accuracy gap to the uncompressed model by 1.41 times on Qwen and 1.44 times on DeepSeek at identical peak memory savings.
Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random
See companion entry.
Predicting Program Exit Code with LLMs and Programming Language Semantics
Code-capable models may succeed at generation while lacking real command of programming-language semantics, and it is unclear whether they apply formal rules even when those rules are supplied in the prompt. The Program Executability Prediction task (PrEx) asks a model to judge whether a program is semantically valid given its syntax and operational semantics, and to name the violated rule when it is not; the accompanying dataset is built by systematically transforming valid programs into invalid ones across human-written, model-translated, and fuzzer-generated splits. Under two semantic formalisms and two semantic shifts, open-source coding models lean on pre-training priors instead of applying the given rules, performing especially badly when the semantics are modified and degrading further as programs grow more complex.
Enoki: Efficient Multi-Level Hallucination Detection
Hallucination detectors generally work at one granularity — claim-level checks give interpretable factual units, span-level checks localize the offending text — and combining them costs extra decomposition, verification, and claim-to-span alignment. Enoki uses open information extraction to pull text-anchored relational facts, verifies each against evidence, and projects unsupported facts back onto spans, so one shared representation serves both levels without any separate alignment step. Extraction can run in language-model, encoder, or rule-based regimes behind a common interface to trade accuracy against cost; the system stays competitive with strong claim-level pipelines at lower resource use and leads on fine-grained span- and entity-level localization, alongside a released dual-granularity dataset called EnokiQA.
Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation Reranking
Reranking methods such as Minimum Bayes Risk (MBR) decoding and Quality Estimation (QE) reranking improve neural machine translation output but require generating and scoring many candidates, and prior acceleration work targeted only the reranking step of MBR. Quit (Quantifying Uncertainty for Incremental Termination) treats candidate generation as a sequential decision under uncertainty, generating and reranking candidates incrementally and stopping once the top estimated quality score stabilizes. Across three translation models and 19 language pairs, it delivers end-to-end speedups of 1.47–2.66× for MBR and 3.43–4.12× for QE reranking while keeping quality inside prespecified equivalence margins.
Topological Steering
Activation- and feature-space steering of LLM behavior operates on local geometry, which makes it sensitive to outliers, noise, and distribution shift. Topological Steering instead represents activation spaces through persistence diagrams borrowed from Topological Data Analysis (TDA), which capture global structure, and uses that representation to drive behavioral interventions. The authors report that the method consistently changes model behavior across multiple model families and sizes.
Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs
Edge deployment decisions are usually made on one metric at a time, so the authors combine accuracy, latency, memory, energy, and safety into a Holistic Sustainability Score organized around economic, environmental, and social pillars. Thirty measured configurations cover five natively small models at BF16 plus five larger models at BF16, INT8, NF4 4-bit, GPTQ 4-bit, and GGUF Q4, evaluated on five zero-shot benchmarks, GPU energy and throughput measurements, and attack success rate on harmful prompts. Qwen3-30B-A3B at GGUF Q4 ranks first overall at 93.38, ahead of Mistral-Small-24B at GGUF Q4, with Phi-4-mini the top-ranked small model, so the assumption that natively small models are always the more sustainable edge choice does not hold universally. The authors stress that quantization behaves as a systems-level choice rather than a smooth precision-efficiency trade-off, and that the score is relative to its comparison pool.
SCoNE: Selective Context-aware Neuron Editing for Robust Retrieval-Augmented Generation
Retrieval-Augmented Generation degrades when retrieved documents mix informative and irrelevant context, since the model gets distracted and hallucinates. SCoNE is a training-free model edit that locates feed-forward neurons scoring high on both attribution and cross-input variability — the ones treated as context-aware — and selectively strengthens them, using only a small set of mining samples, no fine-tuning, and no added inference cost. Across several knowledge-intensive question answering benchmarks and two LLM backbones it consistently outperforms competing noise-robustness baselines, with code released.
Value Over Language Model: Detecting Original Contribution in Writing
Detectors of machine-written text measure how much of a document's surface prose came from a model, not how much of its information content or ideas the human actually supplied. The VOLM (Value Over Language Model) framework never scores surface text at all: it extracts a document's content at increasing levels of granularity, has a language model reconstruct the document from each partial representation, and compares those reconstructions against ones generated from the task description alone, giving a contribution score relative to a replacement-level document. Across news articles, ICLR peer reviews, and argumentative essays it separates human-authored documents from matched model-generated ones while staying largely invariant to content-preserving transformations, including LLM rewriting and round-trip translation. Tightening the content extractor shrinks the residual gap between model-generated and humanized text, which the authors read as evidence that content must be disentangled from style.
Online Self-Weighted Fine-Tuning
Supervised fine-tuning (SFT) weights every expert demonstration equally regardless of whether the model already handles that query, while reinforcement learning adapts update strength but needs far more sampling and can destabilize on hard tasks. Online Self-Weighted Fine-Tuning (OSW-FT) keeps the gradient direction anchored to the expert trajectory but rescales the SFT loss per query by a success rate estimated from a handful of inference-only rollouts, an estimator the authors show is unbiased for the surrogate update at any finite rollout count and connect to SFT and RL through variance-reduction arguments. Across Qwen3 models from 0.6B to 4B on benchmarks including AIME, the method consistently beats SFT on small and medium models using only 2 online rollouts per query.
Can Large Language Models Forecast What Researchers Study Next?
Judging generated research ideas for novelty or feasibility at the moment they are produced says nothing about whether they anticipate what a field actually does next. IdeaForecastBench poses that forecasting task directly: given a community's literature up to a cutoff, a system ranks up to five ideas, which are scored against papers that appeared afterwards, over 624 rolling episodes across 52 topics with a fixed retrieve-then-judge protocol and two separately reported judges. Comparing five history-compression strategies over GPT-4.1, Qwen2.5-7B/14B, and Qwen3.5-9B plus a learned Mode-Decomposition Forecaster, summarizing the history beats feeding it directly on Hit@5 and Precision@5 for all four backbones, with Qwen2.5 scoring above GPT-4.1 by producing broader forecasts; threshold and judge diagnostics temper how much realization should be read as precise anticipation.
How Do Language Models Choose Between Context and Memory?
When retrieved context contradicts what a model learned in its weights, activation directions can steer which source wins — but steering along a direction does not show the unedited model uses it, or that it transfers. The authors estimate authority directions from agreement prompts where context and parametric knowledge concur, then swap naturally occurring coordinates along those directions between matched prompts instructing the model to favour one source or the other. Across Qwen, Llama, and OLMo, the swap reproduces 30-68% of the authority-induced shift in source choice while matched controls reproduce almost none; directions learned on one task close only 9% of the authority gap on another versus 57% for the locally learned direction, suggesting the computation is task-dependent rather than a reusable global feature.
Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning
On in-context learning tasks where a long novel context defines rules, knowledge, and output schema and grading checks every detail, strong open-weights models pass only 12-16% of tasks because one missed rule sinks the whole response. The authors argue the read-and-reason paradigm is structurally to blame — extraction, planning, generation, and self-verification all crammed into one forward pass — and propose the Context Compilation Architecture (CCA), which compiles prose context once into a typed intermediate representation with fixed slots for must-do, must-not, and conditional rules, output spec, available tools, and data profile, then runs executable verifiers and a violation-gated correction loop. Across 1,899 tasks in CL-bench and four open base models it beats vanilla prompting and both long-context baselines (ReadAgent-P, Ctx2Skill) on every model, lifting Kimi K2.5 from 15.4% to 21.4% with gains concentrated on rule-dense categories.
Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation
Parameter-efficient fine-tuning is normally debated in terms of how many parameters to train, but under a severe budget where trainable factors cannot repair a bad subspace, where those coefficients sit matters just as much. The authors study this with frozen-core adaptation — a calibration pass fixes left and right bases per weight matrix and only an r x r core is trained — and propose FCCA, which estimates the signed input-error cross-covariance, whitens it with diagonal Fisher moments, truncates in that local metric, maps back, and applies thin QR for stable core coordinates. Comparing eight basis constructors on 11 tasks, four model settings, and three seeds, FCCA leads at all three Qwen scales (83.0 macro-average on Qwen2.5-3B, 2.3 points above the next matched-budget method) and lands within 0.32 points of LoRA while training 36.9K parameters instead of roughly 7.4 million; ablations attribute 2.7-17.2 points to whitening and find QR necessary for stable optimization.
Instella-MoE Technical Report
Instella-MoE is a fully open Mixture-of-Experts (MoE) language model with 16 billion total and 2.8 billion active parameters per token, trained from scratch entirely on AMD Instinct MI300X and MI325X GPUs. The architecture combines sparse expert routing with Gated Multi-head Latent Attention and FarSkip-Collective connectivity, and the training pipeline runs from pre-training through mid-training, long-context extension, supervised fine-tuning with feedback-driven data curation, direct preference optimization, and reinforcement learning with multi-teacher on-policy distillation. It averages 76.7 across standard pre-training benchmarks, ahead of other fully open models such as OLMo-3-7B and OLMoE-1B-7B while staying competitive with open-weight models like Moonlight-16B-A3B and Qwen3.5-4B, and the post-trained Think checkpoint averages 73.2 across instruction-following, reasoning, math, coding, and chat. Weights, training configurations, data mixtures, and training code are all released.
SFAD: Speculative Factuality-Aware Decoding
Keeping generated text faithful to supplied context is usually bought with extra compute: contrastive decoding needs two forward passes per step, and post-training alignment needs substantial reinforcement learning. SFAD (Speculative Factuality-Aware Decoding) folds faithfulness into speculative decoding instead, training a context-faithful draft model by Direct Preference Optimization on ConFide, a preference dataset built from fine-grained atomic perturbations. At inference an Epistemic Friction score quantifies distributional tension between draft and target weighted by draft certainty, and when it crosses a threshold an asymmetric residual-based logit injection steers the target distribution; otherwise normal speculation continues. The reported result is improved faithfulness alongside a 2.48x speedup rather than the slowdown contrastive methods incur.
Towards a Reliable and Practical Eval Pipeline
Teams shipping LLM-based software increasingly gate releases on "evals", but published work tends to address single aspects of eval reliability rather than what a production pipeline actually needs. The proposed end-to-end pipeline pairs automated creation of eval checklists with a learned aggregation step over the checklist responses, improving both agreement across LLM judges and accuracy against human judgments. It also surfaces self-consistency, explanations, and prediction uncertainty, with empirical results supporting the design.
Replacing Training with Memory: Listwise Selection for Text-to-SQL
Text-to-SQL pipelines that generate many candidate queries and pick one usually rely on a listwise selector that must be fine-tuned, which is expensive. MaP-SQL replaces both fine-tuning objectives with inference-time machinery: reusable structured memories distilled from training data encode how natural language maps to schema elements, SQL operations, and expected outputs, serving as explicit criteria for comparing candidates, while rankings are aggregated across multiple input permutations to cancel positional bias, with execution results and pointwise scoring keeping the comparison count down. On BIRD-dev with the same candidate sets, it beats the prior selector-based state of the art R^3-SQL by 2.02 execution accuracy points while using 2.92x fewer tokens.
CacheBridge: Efficient Cross-Model KV Cache Transfer
When several large language models share context in a multi-model system, the receiving model normally has to re-prefill the shared prefix because key-value (KV) caches are model-specific. A recent training-free approach, Full-Head Mapping, fits a closed-form affine mapper between source and target caches, but maps every target head from every source head, making it fragile to architectural differences and expensive to store and apply. CacheBridge restricts each target head to a single matched source head, weights reconstruction errors by causal attention sensitivity, and builds the mapper with a fused GPU kernel that avoids materializing full observation tensors. It recovers two Ministral 3 transfer directions where Full-Head Mapping loses substantial accuracy, keeps 99.83% mean target retention on Qwen3, and on the Qwen3 14B-to-32B transfer cuts mapper storage by 8x and 500-sequence construction time from 92.63 to 8.63 seconds while matching the baseline with a tenth of the calibration data.
Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO
Language models often ignore prompt evidence that conflicts with memorized knowledge, and post-training can improve this, but it is unclear whether the gains build new machinery or amplify what the base model already has. The authors compare nine post-training arms spanning GRPO, supervised fine-tuning (SFT), and DPO from a single starting checkpoint, extended across scales and families, and estimate a grounding direction from that checkpoint before any training. Five GRPO variants yield small grounding gains even as the rewarded metric improves, conflict-SFT helps moderately, and DPO pushes grounding near ceiling on its matched distribution, yet both SFT and DPO use largely the same causal attention-head set as the starting model. Subtracting the pre-existing direction suppresses both gains, adding it to the base model recovers 35% of DPO's gain at a dose passing all side-effect checks, and a supervised warm start leaves GRPO with essentially no further grounding to add, indicating the gains largely depend on machinery already present.
PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition
Mixture-of-Experts (MoE) inference frameworks treat each expert as an atomic execution unit, which fixes the optimization boundary too early and ignores computational redundancy inside experts. PCoMoE reframes MoE inference as fine-grained path composition, introducing a path-level formulation of expert computation, a compatibility-aware layer-wise pruning strategy that suppresses low-value path combinations, and a hardware-friendly execution engine that exploits reusable sub-expert structures under bounded overhead. The authors report up to a 1.31x end-to-end inference speedup alongside a 10% accuracy improvement, with code released.
Lagged Coupling: Internal Representations Become Readable Before They Become Causal
Probe accuracy is often taken as evidence that a model's internal representation can be used for steering, and the authors test that assumption across the full Pythia suite (160M to 12B, eight checkpoints, four task families) plus a pre-registered OLMo-2 replication. A linear probe reads the target variable from the residual stream at AUROC of at least 0.990 from step 1,000 at every scale, yet steering along that same direction is null-equivalent in 43 of 48 model-checkpoint cells, and the lag does not shrink with scale. They name this structure lagged coupling and decompose it into three tracks (internal readability, behavioral readability, and causal efficacy), finding that representation headroom grows up to 57x with training while causal write-in stays under 0.11% of it. Two pre-registered single-onset hypotheses resolve as indeterminate, and the authors caution against inferring steerability from probe accuracy.
OUTLETS: Output-Length Prediction from Speculative Decoding Backbones
Heavy-tailed output lengths in Large Language Model (LLM) serving complicate resource provisioning and scheduling, and existing length predictors either add latency through external proxy models or rely on shallow probes of current model state. OUTLETS observes that the latent representations produced by the draft decoder in speculative decoding frameworks such as EAGLE-3 already encode signals predictive of generation length, and attaches a lightweight regression head to that backbone to turn it into a trajectory-aware length predictor at almost no added cost when draft states are computed anyway. It achieves lower mean absolute error than the evaluated methods, and under saturated disaggregated serving its predictions let standard scheduling policies prioritize short requests and balance load across decoding instances, cutting short-request P99 latency by 34.8%.
Post-hoc Alignment of LLM-judges to Human Judgment Distribution
LLM-as-a-judge (LLMaJ) evaluations are typically scored against aggregated ground-truth labels, discarding the Human Label Variation (HLV) that a distribution of annotator judgments carries. Across five datasets the authors find LLMs approach human-level accuracy at predicting a single hard label but perform poorly at predicting the soft-label Human Judgment Distribution (HJD). NAPHA (eNtropy-Aware Post-Hoc Alignment) is a lightweight fix that assigns each instance to a discrete entropy class and routes it to a specialized trained alignment model that maps the LLM's distribution onto the human one; it consistently improves soft-label prediction across base models and datasets, with the largest gains on high-entropy instances, and oracle experiments show better entropy-class prediction would raise its effectiveness further.
Beyond Magnitude: Contrastive Routing for Modular Mixture-of-Experts
Routing in Mixture-of-Experts (MoE) layers operates on token representations dominated by structure shared across all tokens, which limits how much experts can specialize. The Contrastive Routing Mechanism (CoRM) contrasts each token against an Exponential Moving Average of the layer's hidden states and scores each expert by the gap between its affinity for the token and its affinity for that shared reference state, through a distinct per-expert projection; the authors show this concentrates the routing signal into a low-dimensional, highly separable subspace with boundaries that align more closely with linguistic structure than standard Top-k routing. On nine zero-shot reasoning benchmarks CoRM improves average accuracy by +0.67 to +1.69 points (Top-1) and +1.38 to +1.77 points (Top-2) over Top-k MoE baselines, at a cost of 2.9% more parameters and 2.6% more FLOPs per token.
Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation
When a hint turns a failing generated program into a passing one, it is unclear whether the hint supplied missing information or merely steered the model toward a solution it could already reach. The authors test this with executable evaluation on HumanEval+ and MBPP+ using Qwen2.5-3B-Instruct and Phi-3.5-mini, comparing adaptive relevant hints, an unrelated hint, and plain repeated sampling without hints, and they add mechanistic probes of a hint-related activation direction. Relevant hints rescue 36 of 79 selected Qwen failures, but eight unhinted samples recover 31 of those 36, and the same pattern holds for Phi, where unhinted sampling recovers 36 of 42 relevant-hint rescues. Persistently adding the shared hint direction yields 14 rescues and 18 regressions with no detectable net gain, so the evidence does not establish task-general capability transfer, though the authors note that differing attempt budgets across conditions prevent isolating a purely semantic effect.
Does task decomposition improve automatic NLG evaluation?
The LLM-as-a-judge (LLMaJ) framework offers cheap, reproducible, reference-free evaluation of Natural Language Generation (NLG), and prior work has tried to improve it by decomposing evaluation into simpler sub-tasks. This study systematically compares LLMaJ methods with and without decomposition across multiple NLG datasets against a fair baseline that does not decompose. There is no evidence that task decomposition improves performance; previously reported gains stem from using human labels as training data rather than from decomposition itself. When human labels are available, LLMaJ without decomposition performs comparably to human annotators.
Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training
Large language models show a modular internal organization resembling functional networks in the human brain, but prior work has only characterized finished models rather than how that organization forms. The authors train Pythia-410M from scratch in two trajectories (bf16 and fp32) and run attribution patching at every step, alongside probes of gradient norms, effective updates, weight norms, and first-order loss decomposition across 14 tasks in four cognitive domains. The modular map is pre-carved: before any learning, the dominant task pair already overlaps at about 3.6 times the task-independent attribution baseline, and the partition then locks in through two sharp jumps whose amplitudes do not track the learning-rate schedule, accompanied by gradient-level relative deprivation in which winning tasks receive 2.25 to 2.73 times the loser's gradient supply without this propagating to updates or weights. Deviation from the baseline substrate appears only in the domain being learned, and the authors pre-register a scale-threshold hypothesis for ongoing 2.8B experiments.
LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs
Flagship language models score above 90% on benchmarks like MMLU, yet fixed question sets test only what experimenters thought to ask. LLMPEDIA recursively materializes about 1.3M encyclopedia articles from the parametric memory of GPT-5-mini, DeepSeek-V3.2, and Llama-3.3-70B without retrieval, then audits a stratified sample of atomic claims against Wikipedia and a curated web stack, labeling each claim supported, refuted, or insufficient. On a uniform random sample only 68.4% of claims are true, more than 21 percentage points below MMLU scores, and 30.5% are insufficient, meaning neither benchmarks nor the world's largest encyclopedia can adjudicate them, whether long-tail knowledge or plausible hallucination. The result is a live, open encyclopedia offering link-traversal exploration, claim-level factuality, cross-model and political-persona comparison, and guided topic drill-down, with every page, claim, and verdict at a stable URL.
FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue
Banking assistants that interact with the same customer repeatedly must keep complete, current, and traceable records as life changes surface incidentally in routine requests, but existing benchmarks test question answering or bounded recall rather than exhaustive longitudinal reconstruction. FinLifeBench contains 6,000 eight-turn Korean banking sessions from 20 synthetic customer trajectories and asks models to reconstruct every life-event instance with its first-establishing session and to rebuild a complete 34-path financial state at consecutive checkpoints, with deterministic gold labels for 24 event types. Across eleven LLMs given the full context, event-anchor recall falls from 0.591 at 15 sessions to 0.445 at 300, with errors driven mainly by omitted events, while financial-state reconstruction frequently treats superseded information as current and the best checkpoint accuracy reaches only 0.470. Performance on the two tasks is only weakly correlated, indicating models can localize evidence for events they recover yet still fail to maintain temporally valid records.
Prompt-Robust Language Models: Which Training Strategies Work?
Large language models remain highly sensitive to how a prompt is phrased, and prior work has tackled this either through refined data construction or dedicated robustness objectives. The authors reproduce and compare these strategies under controlled conditions and find that robustness fine-tuning beats standard fine-tuning and in-context learning, yet the gap between the best and worst prompt template remains 40 to 57% of performance. Recent objectives such as CoIN for contrastive alignment and PPCL for consistency regularization often fail to beat the simplest data strategy of training on one template per batch. Diagnostics show the auxiliary objectives only move the quantity they penalize without generalizing, and that per-template gradients conflict in sign on 57 to 64% of parameters, so mixed-template batches force the optimizer to reconcile competing updates instead of finding a shared prompt-agnostic one.
Post-Training Science for Supervised Fine-Tuning
Every supervised fine-tuning (SFT) run rediscovers the same choices from scratch: learning rate, batch size, LoRA versus full fine-tuning, epoch count, optimiser, and data. The authors run a controlled sweep that varies one lever at a time across dense and mixture-of-experts models from the Qwen3 and Llama families, on four real-world customer SFT datasets whose training data was iteratively refined to pass a customer-built evaluation, for both LoRA and full fine-tuning. The study asks how optimal learning rate and batch size shift with scale, family, and data and whether one selection rule transfers; what LoRA trades against full fine-tuning and how rank and alpha bound what an adapter can learn; whether validation loss or loss-landscape flatness faithfully ranks downstream quality; how gains scale with model size and data on a ladder reaching 235B parameters; how many epochs can run before general instruction-following erodes; and whether a geometry-aware optimiser beats AdamW. Each recommendation is paired with a measure of its uncertainty.
mzCache: On-Device LLM Memory Management under Multitasking
Phone users switch apps constantly, so the operating system evicts model weights and key-value cache under memory pressure, forcing an on-device language model to reload from slow storage or recompute its whole cache when the next request arrives. mzCache manages memory for this multitasking regime by splitting model memory into fine-grained shared buffers that support partial eviction and restoration, and by using the unified memory of mobile systems-on-chip to keep GPU inference running while the CPU restores in parallel, with hybrid swap and backward-out eviction policies. Implemented on llama.cpp and deployed as an Android application, it cuts time-to-first-token by 2.1 to 5.5 times relative to storage-backed partial offload.
Probing Factual Knowledge Transfer with Training Data Interventions
Whether multilingual models genuinely carry facts across languages or mostly recall facts seen in the target language is hard to test observationally, so the authors intervene on the training data: starting from an English-pretrained model, they continue pretraining on Persian text with specific facts systematically removed at several levels of granularity. Their SIFT resource contains 500 triples across 20 relations, split by whether the subject is globally prominent or Persian-specific, with natively written Persian cloze templates. Under the strictest removal condition a large majority of English-acquired facts fail to transfer into Persian, sentence-level co-occurrence filtering leaves fact signal intact, and randomly chosen negative candidates inflate apparent transfer by rewarding shallow associative heuristics. Facts about Persian-related entities, far rarer in the English corpus, barely transfer at all.
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
Looped transformers add effective depth by iterating a shared block, but comparing at fixed model size hands the looped variant extra floating-point operations, conflating architecture with compute. SMELT loops the middle half of layers twice in a Mixture-of-Experts transformer while matching per-token FLOPs, total non-embedding parameters, and key-value cache against an unlooped baseline, scaled across four sizes up to 54B non-embedding parameters with a separate Chinchilla-style scaling law fit per architecture. Loss falls faster with compute, saving 6.8 to 18.0 percent of training FLOPs on the compute-optimal frontier, with downstream gains exceeding what validation loss predicts, largest on code, and growing with sequence length and in-context example count. Mechanistic analysis credits the second visit with shrinking the attention sink and redirecting attention toward content-relevant tokens.
Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades
Inference cascades route most queries to a cheap model and escalate a hard tail to a frontier model acting as verifier, and a tempting extension finetunes the cheap student on the verifier's rejections so escalation and cost fall each round. Measuring the loop on real models, the authors find the verifier's blind spot, the share of student errors it wrongly accepts, grows with student capability (0.12 at 0.5B parameters to 0.55 at 32B) and shrinks with verifier capability, so it is worst in the cheap-student, cheap-verifier regime cascades exist to create; a frontier verifier nearly closes it but then escalates 46 percent of hard MATH queries against a 39 percent true error rate. Corrective finetuning on the rejected tail degrades and eventually collapses the small student across every teacher tried, cross-family and same-family alike. Throughout, every metric computed through the verifier reads a flat 3 percent error while true delivered error swings as high as 32 percent, a blindness the authors formalize as a two-population conservation law and validate synthetically.
CHARM: Character Hallucination for Multicultural Role Play Benchmark
Role-playing language models are supposed to hold a character's voice while respecting that character's knowledge limits, but existing hallucination evaluations do not separate failing to notice a boundary from crossing it after noticing. CHARM covers 40 real and fictional characters from five cultural-linguistic regions, validated by native reviewers, probing temporal and cross-universe boundaries with abstention-enabled multiple-choice questions and a two-stage score that splits boundary awareness from boundary compliance. Across six models, hallucination is driven predominantly by compliance failures: the model states that a question lies outside the character's knowledge and then answers it factually anyway. Re-posing the same questions to the character confirms many cases are parametric overrides where the fact is stored but not suppressed, and failure rates vary systematically by cultural region.
Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs
Multilingual models translate well, yet how they transform a representation from one language into another is only partly understood; prior work splits the process into language-independent conceptual content followed by production into language-specific form. Using controlled multilingual datasets that isolate word-order differences, plus causal interventions and probing, the authors show the production stage splits further, with models committing to target-side word order before realizing the target language's surface form. They also identify individual attention heads that respond selectively to syntactic transformations while remaining largely invariant to language identity, placing syntactic commitment as its own stage in the translation pipeline.
Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA
Linear classifiers trained on a model's hidden states can flag factual errors in a single forward pass, implying true and false statements separate along a stable truth direction, but published results disagree on whether that direction survives input shifts because cross-dataset transfer experiments change several variables at once. The authors isolate writing style, medical specialty, and source corpus by rewriting 500 MedQA items into textbook, patient, clinical-note, and colloquial registers, annotating each with a specialty, and grouping them with MedMCQA and MMLU-medical. Probing four open-weight models of 2 to 8B parameters, style costs about 0.10 AUROC and specialty about 0.03, while corpus shift costs up to 0.21 AUROC, roughly twice the register gap. The register result replicates with a second generator and with human-written patient questions, so question format does not explain the corpus break, suggesting the probe signal is partly bound to dataset structure rather than medical knowledge.
How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation
Scoring free-form question answering is a bottleneck because a correct answer takes many surface forms and a wrong one can fail by incompleteness, contradiction, overgeneration, or endorsing a false premise, distinctions that judge-based and similarity-based metrics collapse. The authors define an eight-class ordered taxonomy of semantic correctness that keeps verbose-but-correct answers apart from answers carrying hallucinated content, and release CAP-Correctness with 8.8k examples across widely used question-answering datasets plus CAP-Statements with 11k question-answer-to-statement conversions for natural language inference (NLI) training. Their reference-based metric CAP (Context-Aware Precision) scores question-conditioned statements with bidirectional NLI and outperforms established baselines under a monotonicity protocol that tests whether a metric respects the taxonomy's intended ordering.
Behaviorally Effective LoRA Writes Are Sparse and Structured
Low-rank adaptation fixes the rank of an update but says nothing about which parts of a trained adapter actually carry behavior. The authors warm up an unconstrained adapter, convert its learned write columns into a frozen module-wise orthonormal basis, and continue training inside that constrained parameterization; across 14 exact switches held-out accuracy is unchanged at conversion and reconstructed write matrices differ by at most 0.25 percent relative Frobenius error, while continuing the same checkpoint under different write subspaces leads to different outcomes, marking write geometry as a causal state variable. A no-retraining projection test shows the useful signal stays inside the learned write space and vanishes in random or frozen-activation principal-component controls. On GSM8K, MathQA, and AQuA, per-module top-k continuation peaks at k of 2 or 4 in all twelve seed-level cases, learned top-16 and top-32 subsets beat matched random subsets, and single-direction ablations isolate a few late query, output, and down projection components with outsized behavioral impact.
When Tokenization is Secretly Output Supervision
Tokenization is normally treated as an input preprocessing choice, but in autoregressive models the tokenizer also determines what must be resolved in a single forward pass and therefore what supervision signal the model receives. A controlled numeric-reasoning experiment that decouples input from output tokenization finds that differences in task performance, training dynamics, and model internals are induced by output tokenization and are largely invariant to input tokenization, implying that models with different tokenizers were effectively trained on different tasks rather than differing only in ability. A survey of 120 recent computational-linguistics papers on numeric reasoning finds only about 10% report the numeric tokenization of the models they evaluate, while 69% compare across tokenization regimes without reporting it.
Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search
Optimal hyperparameter scaling laws let practitioners predict good configurations at production scale, but fitting them conventionally means exhaustive grid searches over thousands of training runs. Power-Law Entropy Search (PLES) is a cost-aware acquisition function for multi-fidelity Bayesian optimization that selects, at each iteration, the configuration that most reduces uncertainty in the scaling-law estimate per unit of compute, targeting the law itself rather than a single objective and naturally favoring cheap small-scale experiments. Across synthetic benchmarks, surrogates fitted to real large language model training data, and actual pretraining runs, it converges to accurate scaling laws using less than one-tenth of the compute required by grid search and other baselines.
Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation
Citations carry rhetorical intent — supporting, contrasting, or merely mentioning prior work — and whether models reproduce that intent when assisting scientific writing is untested. A masked-citation task has six popular large language models generate replacement citation sentences, producing a counterfactual corpus directly comparable to human writing over 1,746 top natural language processing conference papers, 63k+ contexts, and 132k+ citations, with an LLM judge classifying intent and a 20-million-edge coauthorship network measuring social distance to cited authors. Models cite significantly less critically than humans, over-cite popular and older papers (a tendency strongest where humans would contrast against recent, niche work), and draw on more socially distant authors than the close collaborators humans favor for supporting citations.
LatentPress: Context Compression Beyond Text and Vision
Compressed context is normally carried as readable text or rendered images that must be decoded, even when the consumer is a language model. LatentPress writes conversation histories and long documents into continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, training only a 4.2M-26.2M-parameter adapter (roughly 0.1% of the decoder) to compress 4-16x with no text reconstruction at inference. On LongMemEval it reaches 0.504 accuracy at 7.70x compression, above the 0.490 obtained from uncompressed evidence, and far above text summaries (0.184) or OCR-based compression (0.426 to 0.312); writing costs 43ms per conversation, about an order of magnitude faster than summarization or OCR, and reading is 5-9x faster than raw context. Transfer holds zero-shot from UltraChat to LongMemEval and onward to unseen LongBench document domains, though 16x compression still trails raw context.
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Logit-based knowledge distillation trains small language models from stronger teachers, but its benefits turn out to depend on the training stage. Forward Kullback-Leibler distillation improves both reasoning and factual recall during pre-training, yet during mid-training — the intermediate self-supervised phase on curated corpora — it keeps delivering reasoning gains while slowing factual recall, which the authors trace to teachers being more confident on procedural than knowledge-intensive data while students acquire low-entropy facts early. Switch Distillation uses teacher predictive entropy as a routing signal, distilling only where the teacher is confident and falling back to cross-entropy elsewhere; relative to standard next-token prediction it achieves 1.61-1.71x the reasoning performance while preserving 96.7-96.8% of factual recall, and the advantage survives post-training.
Can LLMs Discover Scientific Laws in Real and Parallel Worlds?
Whether large language models can genuinely discover scientific laws is hard to judge because existing evaluations either use simplified synthetic setups or reuse published targets the models may already have memorized. SciLaws-Bench draws 118 problems from 381 papers, covering 291 candidate laws and roughly 8 million real data points across six disciplines, and poses each in two settings: SciLaws-Real, where models propose laws from fixed real observations and are judged on held-out predictive fit and literature-derived scientific validity, and SciLaws-Parallel, where models actively query residual-calibrated simulated worlds to recover a newly synthesized hidden law. Predictive fit can diverge from scientific validity, memorization determines whether models reproduce or move beyond published formulas, and a best-of-N study reveals a selection bottleneck.
Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories
Embedding retrieval is tested where surface form and meaning are deliberately pulled apart, retrieving items that share underlying structure but not wording, on competition mathematics (MathNet-Retrieve, 500 queries over a 117,088-item corpus) and ALFWorld-derived embodied-agent trajectories (118 queries, 336 trajectories). In mathematics, strict Hit@1 at the heaviest disguise tier is 0.0% for both production embedders, even though the correct item almost always sits in the top 10 and the winner is more lexically similar to the query in 95.2 to 99.8% of misses; in trajectories the same models fall to chance or below once the gold item must differ in object and receptacle. A lexical reranker hurts in mathematics but helps in trajectories, which the authors use as a diagnostic for whether a benchmark's surface variation is adversarial or incidental, while an LLM reranker recovers 5 to 63% of the gap in mathematics and 43 to 76% in trajectories, with part of the mathematics gain traced to memorization of well-known competitions. A paired downstream experiment found oracle retrieval indistinguishable from adversarially bad retrieval because the solver's zero-shot accuracy was largely a truncation proxy, leaving no headroom for retrieval to matter.
From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification
Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, and the usual fix of retrieving the top-K candidate labels by embedding similarity narrows the choice without helping the model tell near-duplicates apart. The proposed framework identifies which label pairs the model confuses, expands the candidate set to include those confusable labels, and generates targeted rules that distinguish similar candidates, all without fine-tuning. On WOS, Flipkart, and LEDGAR, Macro F1 improves by up to 10.0 percentage points over retrieval baselines, and the generated rules transfer to smaller 2B to 20B models, which gain up to 11.5 points.
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
Data-residency rules push enterprises to self-host large language models, and adopting newer models without retiring their predecessors fragments a finite GPU pool across a growing serving fleet. The authors consolidate traffic from over 200 internal applications onto a single model by closing quality gaps found through production error analysis along instruction following, function-calling, and the internal task distribution, tracked with offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimizing all objectives jointly, which caused cross-domain reward interference, they train a separate GRPO expert per axis and merge them with two-stage SLERP, with each expert's reward exposing a distinct failure mode: semantic collapse, over-calling, and verbosity hacking. In non-reasoning mode the model beats a roughly 7x larger baseline on an in-house Arena (69.6 vs 65.8), instruction following (0.85 vs 0.83), and function-calling (0.79 vs 0.77), and it now absorbs 50% of platform traffic, 116 million requests per month, at a fraction of the serving cost.
Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
How to split a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) in LLM post-training is unresolved, with prior work offering only broad trends and no evidence on whether the best ratio transfers across model sizes. Instead of hunting for a single optimal ratio, the authors characterize the near-optimal region, the set of allocations within a given tolerance of peak performance. That region is wide even at 2 to 10% tolerances, widens with model scale, and transfers reliably from small proxy models to large targets, so cheap proxy experiments suffice to pick an allocation without exhaustive large-scale search. The pattern holds across tasks, model families, and both preference-based off-policy and reward-supervised on-policy RL methods, and the authors show how asymmetric annotation costs for SFT versus RL data shift the region.
The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally
Post-training quantization (PTQ) cuts LLM serving cost, but its accuracy damage is uneven and typically tuned per model. Using causal mixed-precision intervention as ground truth, raising each layer to 8-bit in turn and measuring recovered accuracy across 9 open-weight models from 4 architecture families, the authors test whether damage lives in task circuits, where the model computes, or in weight statistics, and find that none of these predicts which layers benefit from restored precision. Recovery is instead diffuse: for 8 of 9 models, recovering 75% of the gap takes roughly half the layers, with Qwen3-8B the lone sharply concentrated exception. At a matched precision budget, spending it globally on finer quantization granularity beats locally repairing the most recoverable layers by 21 to 52 points for every group-128-compatible model, and since 8-bit is near-lossless across RTN, GPTQ, and AWQ, the authors conclude that cheap correlates of quantization damage must be checked by causal intervention before guiding precision allocation.
Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation
Repository-level code generation must produce code consistent with a codebase too large to fit in a model's context, so most systems rely on retrieval-augmented generation (RAG) that supplies repository context as task-level support. ACToR instead identifies critical tokens, the few decisive positions during autoregressive decoding where an error sends the rest of the output down a wrong semantic path, and triggers targeted retrieval on demand at exactly those positions, aided by a position-aware weighting scheme that makes dense retrievers prioritize context most informative for generation. On RepoExec and CoderEval it consistently beats state-of-the-art methods, with relative improvements of 8.4% and 15.4% respectively, and an accompanying analysis quantifies how heavily major generation failures concentrate at these critical tokens.
Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation
Large language model judges are widely deployed to score natural language generation (NLG) quality and to supply automated training signals, yet the internal procedure by which they assign a rating is poorly understood. The study probes this mechanistically with an eight-attack perturbation taxonomy over the Readability and Adequacy dimensions, a pipeline producing paired clean and corrupted summaries with controlled error intensity and token-level modification maps, and a battery of causal tracing, logit-lens projection, and attention-head knockout applied to Themis (Llama-3-8B) and Prometheus (Mistral-7B). Both evaluators implement a two-stage pipeline: below layer 15, attention performs local error comparison and routes the result to the final input position, while above it the MLP cascade integrates the signal and writes the rating, with the decision crystallizing sharply at layer 26 on Themis and layer 25 on Prometheus. A base Llama-3-8B control reproduces the routing and crystallization but not the stage separation, isolating two effects that fine-tuning specifically installs, suppression of early MLP contributions at the last position and a two-layer earlier crystallization, which indicates fine-tuning sculpts an existing substrate rather than building the pipeline from scratch.
19 more specialized papers
- RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks Xingran Chen, Rohit Bhagat, Ghadir Ayache et al.
- LOOMSUM:Weaving Quantitative and Narrative Evidence for Faithful Long Text-Table Summarization Meng Zhou, Wenhao You, Wei Yuan
- Emotional Labor Strategy Preferences in LLM Personas Mohammad Saim, Tianyu Jiang
- Two locked tests of phase-structure features for transition prediction Abraham Chachamovits
- Location-Aware Language Models via Secondary Embeddings Gokul Srinivasagan, Munir Georges
- Breaking the Structural Identity: Personalized Federated LoRA Fine-tuning under Rank Heterogeneity Lei Wang, Jieming Bian, Letian Zhang et al.
- S^3martCirc: Self-supervised Smart Circuit Discovery Wendy Zheng, Yinhan He, Liang Wu et al.
- RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation Zhongru Chen, Yuan Wu, Yi Chang
- VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences Yiwen Jiang, Yang Deng, Stephanie Fong et al.
- Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages Michele Ciletti
- From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion Satoshi Hayakawa
- StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions Chao Gao, Haijiang Liu, Qiyuan Li et al.
- EDRAC: Benchmarking Arabic Dialect Reading Comprehension Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn et al.
- Overfitting Mitigation via Singular Value Decomposition in Minimum Bayes Risk Decoding Riza Setiawan Soetedjo, Yusuke Sakai, Hidetaka Kamigaito et al.
- Subword Segmental BabyLMs: Learning to Tokenise for Sample-Efficient Pretraining Francois Meyer
- H2Table: Hierarchical Hypergraph-Enhanced Large Language Models for Complex Table Reasoning Jia Ling, Yangfan Wang, Chen Tang et al.
- Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation Thibaut Thonet, Jos Rozen, Laurent Besacier
- Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models Tian Fang, Ga\"el Guibon, Davide Buscaldi
- Polish ModernBERT: The Long and Short of Polish Language Understanding Micha{\l} Pere{\l}kiewicz, S{\l}awomir Dadas, Rafa{\l} Po\'swiata et al.
Applications 98
Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models
Scams aimed at older adults unfold over several conversational turns, escalating from casual contact through trust building and urgency to a request for money or credentials, so risk has to be re-estimated as the dialogue grows. The proposed framework accumulates turns incrementally and re-scores risk at each step, trained on a purpose-built multi-turn dataset of investment, charity, and tech-support scams annotated at every cumulative stage with a risk level, continuous score, rationale, and safety recommendation. Four compact models — Phi-4, LLaMA-3.2, DeepSeek-R1, and Qwen3 — were fine-tuned under a shared recipe, with Phi-4 and LLaMA-3.2 giving the strongest turn-aware risk estimates relative to their parameter count, supporting on-device deployment where conversations never leave the phone.
Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation
LLM rerankers for conversational recommendation are compared against collaborative-filtering and sequential baselines inside a shared retrieve-then-rerank pipeline on the ReDial movie benchmark, while varying candidate-pool size, first-stage retriever, and decoding temperature. With a shared semantic top-250 pool and strict candidate-aware scoring the best proprietary reranker reaches NDCG@10 of 0.1497 against 0.0939 for the strongest non-LLM baseline, yet the same model scores 0.2925 under unconstrained zero-shot generation, and no open-weight model beats a tuned shallow autoencoder. Switching from semantic to collaborative-filtering candidates raises NDCG@10 by more than 50% for the strongest rerankers, and higher temperature mainly changes list stability rather than mean accuracy for strong models, leading the authors to argue candidate generation, pool size, scoring policy, and decoding configuration belong in required reporting rather than in implementation footnotes.
Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization
Steering molecular generators toward desired properties usually relies on policy-gradient reinforcement learning, which needs a trajectory log-probability whose form depends on the specific architecture and generation procedure, so optimizers do not port across model families. EW-SFT (Elite-Weighted Supervised Fine-tuning) instead uses the reward only to select an elite set of high-scoring molecules, then updates the model with its own pretraining loss on that set; ablations indicate the reward signal flows mainly through elite selection rather than continuous weighting. Because the rule needs only scored molecules and the model's native loss, the same optimizer works across autoregressive, masked-diffusion, and discrete-flow generators and across de novo, motif-extension, and linker-design tasks, outperforming each model's native optimizer under a fixed budget of 3D shape alignment oracle calls on two kinase references.
Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models
Crisis helplines assess suicide risk through structured interviews that are slow and depend on operator training, and almost no prior work covers Arabic-language calls or works within real helpline privacy constraints. Using de-identified transcripts from Lebanon's National Lifeline — transcribed on site with a Levantine Arabic speech recognition model and scrubbed locally by an Arabic named-entity recognition model — the authors fine-tuned five instruction-tuned language models and six transformer encoder baselines on both the Arabic transcripts and machine-translated English versions, labeling calls with two binary outcomes derived from the Columbia Suicide Severity Rating Scale. Across 383 calls, the best English model reached a macro-F1 of 85.00 and ROC-AUC of 92.59 on high-risk classification, catching 88.9% of high-risk calls, with the best Arabic model close behind at 81.19 macro-F1; lower-severity ideation proved considerably harder in both languages.
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
Research on AI companionship is limited by scarce and unreliable human-chatbot interaction data. CompanionSim scales a small amount of real data into 2,240 simulated multi-turn conversations covering 16 chatbot behaviors across seven use cases, which human annotators then rated alongside real conversations in two studies (a U.S.-representative sample of 628, and 3,646 participants across the U.S., U.K., India, and Nigeria). Companionship behaviors such as validation reduced likability, humanlikeness, and trust rather than increasing them, with the effect strongest among women and older participants.
NSIDDx: A Design Framework for Neuro-Symbolic, Practitioner-First Differential Diagnosis in Low-Resource Settings
Diagnostic pipelines built on a language model plus retrieval over rare-disease knowledge score well on benchmarks, yet evaluation on uncommon presentations across two cohorts showed they emit confident outputs that clinicians often cannot verify and that resist interrogation. NSIDDx responds with a design framework treating the clinician as an active reasoning agent rather than a recipient of answers, instantiated as a neuro-symbolic pipeline with ternary symptom encoding, contradiction detection, audit strings, and practitioner override that runs offline on consumer hardware. The authors distill five design principles for clinician-in-the-loop clinical language processing and call for prospective studies, positioning the work as a framework proposal rather than a validated system.
Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching
As service marketplaces shift from fixed request forms to language-model matching that infers intent from free text, the provider-side attribute taxonomy underlying matching, search, and pricing has to be rebuilt in a form providers still understand. The described autoresearch loop generates that taxonomy one occupation at a time through iterative propose-evaluate-keep cycles, scoring each candidate tag set with a recalibrated six-rubric model-as-judge and applying weighted penalties from a seven-critic persona panel with no hard vetoes. A separate parity-mapping stage infers which provider attribute each legacy form question was meant to measure and maps it onto the generated tags, giving both a coverage signal and a human quality-assurance interface. The system has run in production at a major U.S. consumer services marketplace since April 2026 across 132 occupations.
Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries
As health information seeking moves from ranked link lists to conversational answers, the job of judging sources shifts from the user to the platform, yet little is known about what these systems actually cite. An audit of ChatGPT, Perplexity, and Google AI Overview on twenty English mental health questions under two prompt conditions, plus a three-question subset translated into six additional languages, recorded 15,942 citations across 1,140 responses and 1,713 unique domains, each classified by a validated nine-category typology. Citations were highly concentrated: the ten most-cited domains accounted for 43.6% of English citations, with government, commercial health, and academic sources each near 22%, and explicitly asking for sources changed the mix only modestly. Non-English queries returned fewer citations and were routed to language-appropriate resources significantly less often; the typology, classifier, and annotated corpus are released for reuse.
Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts
Real-time agent-assist tools in contact centers must decide, for each of many predefined topics, whether a live customer utterance is relevant, working from ASR (Automatic Speech Recognition) transcripts of spontaneous phone calls that are unclear, repetitive, and largely unpunctuated. The authors curate a human-annotated dataset of topic-utterance judgments from real call-center transcripts and compare a regex baseline, zero-shot sentence embedding encoders, and Gemini-based matchers, crossed with two ways of describing a topic: keyphrases versus natural language descriptions. Lightweight LLM matchers paired with natural language topic descriptions outperform both embedding and regex approaches, indicating that how a topic is expressed matters alongside the matcher itself.
Hidden relationships in a document-derived property graph: top-k chunk embeddings and inverse-distance weighting over a dynamically evolving ontology
Knowledge graphs that large language models extract from text capture only explicitly stated facts, leaving semantically related entities disconnected across documents. The proposed additive second pass leaves those facts untouched: each document is chunked and embedded once, top-k nearest-neighbour queries over existing chunks yield candidate node pairs via entity membership maps, and pairs are scored with Shepard inverse-distance weighting over a rescaled chord distance, avoiding the threshold collapse of affine cosine scoring behind a k-NN gate. Un-gated per-pair accumulators form a commutative monoid, making the pipeline strictly order-independent and incrementally scalable without recomputing earlier documents. Implemented across FalkorDB, Kinetica, ArangoDB, and Neo4j, 768- and 240-dimensional embeddings retain 92% and 72% edge fidelity against a 3072-dimensional baseline while the top-k formulation runs 25x faster.
Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations
Managers heading into hard conversations with employees need to practice speaking aloud, which text chatbots cannot offer despite being the scalable option. Conversation Coach is a voice-first rehearsal system built around three requirements: low-latency interaction with strong language understanding, configurable bot personalities that simulate different employee types, and personalized feedback on content and policy compliance. Comparing an end-to-end speech-to-speech model against a cascade of automatic speech recognition, a large language model, and text-to-speech, the end-to-end route delivered 3x lower median (P50) latency with native barge-in at an estimated 8x lower cost, while the cascade reasoned well enough to matter for coaching quality and was the architecture actually deployed. In production it was used by more than 40,000 managers over six months, with adoption concentrated on genuinely difficult conversations.
EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models
Locally run small language models are appealing for scientific question answering because of privacy and deployment stability, but they face small literature collections, fragmented evidence, short context windows, and weaker reasoning. EGT-KG (Evidence-Grounded Typed Knowledge Graph) is a retrieval framework built to work under those constraints, compared against plain retrieval-augmented generation in two variants, one with an automatically generated relation schema and one with an expert-defined schema. Scored on a six-dimensional rubric covering soundness, correctness, completeness, conciseness, relevance, and fluency over a biopolymer-bound soil composite literature benchmark, EGT-KG beats vanilla RAG in most settings, with the largest gain on llama3:8b at a final score of 70.37, up 14.67%.
Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing
A study of 50 participants learning nuclear safety protocols compared three AI tutoring designs: an unrestricted ChatGPT-style chatbot, a Socratic bot that gives hints but withholds answers, and a non-conversational tutor that adapts difficulty from Muse headband EEG signals measuring cognitive engagement. The unrestricted chatbot produced the highest immediate post-test learning gains (p < .03, d > 0.80), while the adaptive EEG-driven condition generated the highest measured brain engagement (p = .018). Clustering of interaction logs showed unrestricted-mode users mostly retrieved answers directly, whereas Socratic-mode users started reasoning through hints and then progressively disengaged. The authors argue the unrestricted chatbot's advantage reflects the immediate timing of the test rather than deeper learning.
CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN
Telecom small language models (SLMs) embedded in AI-native 6G radio access networks typically justify decisions after the fact, and attempts to train pre-hoc reasoning traces with Group Relative Policy Optimization (GRPO) hit a cold-start barrier where models learn either the output format or the correct label but not both. CRAFT (Cold-start Reasoning Alignment via Fine-Tuning) sidesteps this by autonomously generating a verified dataset of input–trace–label triplets and fine-tuning with low-rank adaptation (LoRA). On the TRACTOR and IC xApp datasets it reaches up to 86.5% accuracy and 94.6% F1 with zero parse failures, while direct GRPO and supervised-fine-tuning-then-GRPO stay below 53.5% F1, and it uses 59% less energy; CRAFT-initialized policies also remain stable under subsequent GRPO training with varied reward functions.
TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data
Audio-to-score transcription models are limited by the scarcity of paired audio and notation data, which confines most systems to a single instrumentation. TUTTI drops real scores entirely for pre-training: a symbolic music generation model produces a large multi-instrumentation corpus, which is rendered into audio-score pairs with expressive acoustic variation and used to pre-train a standard Transformer encoder-decoder. Pre-training on synthetic multi-instrumentation data yields stronger representations than single-instrumentation training, and after fine-tuning on real datasets the model sets new state-of-the-art results across audio-to-score baselines and transfers competitively to instruments never seen in training; code and the TuttiCorpus dataset are to be released.
SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification
When a language model rewrites a mathematical optimization problem into another modeling language, checking the result by running a solver is unreliable, since local minima, timeouts, and numerical artifacts can mask genuine semantic divergence. SOVER splits the job in two: the model extracts a variable mapping between formulations, then Z3 formally certifies domain cross-feasibility and objective-order preservation for mixed-integer linear problems while dReal supplies tolerance-aware feasibility, range, and approximate-argmin checks for continuous nonlinear ones. On NLEquiv-150, a new public benchmark of 100 equivalent and 50 deliberately hard non-equivalent nonlinear reformulation pairs, the system classifies 149 of 150 pairs correctly including all 50 hard negatives, with the single failure traced to an incomplete mapping extraction.
AnalysisBank: An Expert Analysis Pattern Library for Financial Report Generation
Automated financial report generation typically plans at the structural level, deciding which topics or sections to include, which tends to produce content that restates rather than analyses. AnalysisBank distils expert reports into a library of Analyses, each binding a data signal to an analytical move and the expert text span it came from, then at inference matches the input's signals against the library and applies the retrieved moves. Distilling 550 expert reports yields a heavy-tailed distribution of 47 to 52 signal types across 13 move types, and on two financial benchmarks with four LLM backbones the approach raises the share of novel, data-grounded insights by 1.7 to 3.7 times over structural-level baselines, with transfer experiments on scientific writing suggesting the analytical-versus-structural distinction is not finance-specific.
Staged Linguistic Seeding: Grounded Query Expansion for Verified-Unit QA in AI Contact Centers
Customer-service question answering in a voice contact center faces latency limits and a high cost for wrong or unsupported answers, so the deployed system answers only from a closed set of human-verified units, returning one verbatim or routing to clarification, abstention, or a human handoff. Coverage comes from staged linguistic seeding, an offline index-enrichment step where a human writes a per-unit grounded slot recipe, gpt-4.1-mini renders it into query variants, and a light human gate filters them, leaving inference as a single retrieval pass with no query-time generation. On held-out variants from two industrial domains, recall@1 rises to 0.881 and 0.930 (gains of 0.27 and 0.34) with improvements across all five retrievers tested, beating doc2query at the same generation budget, and the verified-unit design cuts unsupported content from 7-13% to roughly zero.
Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair
When a code model's first attempt fails a test, fault localization is supposed to help by pointing the repair at the implicated statements, but a targeted edit may simply benefit from being small, and a second sampling call may fix things without using the failure at all. Three arms applied to the same failed candidates — blind whole-solution resampling, spectrum-based localization with suspect-span infilling, and same-length infilling at a random disjoint span — were run across three frozen 26-32B models, three benchmarks, 488 failing candidates, and a separately declared 24B fourth model. Localization was usable on only 9.0% of failures, and among the 177 localizable candidates localized infilling lost decisively to blind resampling at matched attempt counts (3:40, p = 3.0e-9), a result that replicated in a third model family; infilling reproduced the removed span verbatim 48.9% of the time, explaining why extra budget does not help, and the advantage over the random-span placebo held only pooled, not per model.
Conditional Flow Matching for ML-Based Inverse Design Problems
Engineering inverse design is slowed by iterative solvers for problems constrained by partial differential equations and by their sensitivity to where optimization starts, which motivates generative models that propose candidate designs without rerunning the simulator. Conditional flow matching is added to EngiOpt and compared against a conditional diffusion model and a conditional generative adversarial network on the beams2d structural and heatconduction2d thermal tasks from EngiBench, scoring generated designs as warm starts for gradient refinement via cumulative and final optimality gap. Flow matching achieves the lowest measured optimality gaps, maximum mean discrepancy, and volume-fraction deviation on both tasks, and at 16 Euler steps reaches 53.2 samples per second on beams2d, roughly 66 times the throughput of the diffusion baseline at 1000 network evaluations.
A Dataset for Modeling Iterative Problem-Solving
Solving a problem through repeated attempts is a sequential process in which a solver receives feedback and revises, and predicting whether performance improves, plateaus, or regresses matters for understanding both human learners and autonomous agents. CodeInsight captures this at scale with over 3 million submissions from 3,286 undergraduates in two introductory C++ courses over two academic years, including test-case-level outcomes, timestamps, and source code. A benchmark on the dataset compares parametric, sequential, and generative predictors under a shared calibration-and-scoring protocol, and a Recurrent State Space Model (RSSM) adapted to track solver traits through discrete latent variables is the most accurate on three of four courses. An LLM-based predictor that writes full submissions is less accurate but exposes failure modes, and its coding proficiency turns out to be inversely related to predictive performance, making it better understood as a generative solver than a faithful model of student behavior.
ClinTraceBench: Source-Verifiable Longitudinal Clinical Reasoning over EHR-Derived Dialogues
Clinical assistants built on large language models (LLMs) must reason over multi-visit patient histories, but whether compact history representations such as retrieval, summaries, or agentic memory preserve the needed longitudinal signal has not been measured. ClinTraceBench provides 385 verified dialogues derived from MIMIC-IV with event-level provenance and a nine-task taxonomy, used to evaluate eight history strategies, including full context, BGE-M3 dense retrieval, LLM summaries, and the agentic memory systems Mem0 and A-Mem, across DeepSeek-V3, GPT-4o-mini, Haiku 4.5, and Sonnet 4.6 on 6,271 questions. In a controlled injection probe, Mem0, A-Mem, and LLM summarization recover only 0 to 5.3% of injected attribution facts even when the sentence is present before memory construction, and compressed strategies pay an aggregation tax on multi-visit trends and cross-patient comparisons. The gap between no context and full context ranges from about 30 to 63 percentage points depending on backbone, and on the cost-accuracy frontier Haiku 4.5 with full context dominates Sonnet 4.6 at roughly a quarter of the cost.
When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting
Online adaptation can help time-series forecasting on edge devices under distribution drift, but the measured benefit depends heavily on evaluation choices. Using six public multivariate streams including building-sensor and smart-meter data under a leakage-free streaming protocol, the study identifies two sources of comparison bias: the static baseline's warmup budget, which shifts the estimated adaptation benefit by 3.0 to 18.8 percentage points across a 1,000 to 20,000 step range, and comparing SGD with momentum against Adam at a shared default learning rate, which conflates optimizer quality with rate sensitivity. When warmup and per-optimizer learning rates are selected on a held-out pre-drift validation slice, Adam outperforms SGD with momentum in 310 of 360 evaluated cells, though four Adam cells remain below the static baseline. Measurements of adaptation-state memory and A100 per-update latency on PatchTST show several parameter-efficient variants are nondominated on the memory axis, and reported smart-meter gains depend on meter-selection rules.
Athena: Vulnerability-Affected Library Identification via Knowledge Graph Completion
More than half of vulnerability database entries have missing or incorrect affected-library information, and existing automated approaches treat identification as isolated text retrieval, ignoring the relational structure of the databases. Athena models vulnerability databases as a knowledge graph integrating CVEs, libraries, CWE weakness types, CPE products, and software ecosystems, reformulates the problem as knowledge graph completion (KGC) via link prediction, and re-ranks KGC candidates with a fine-tuned LLM augmented with knowledge graph embeddings. On VulLib, Athena improves average F1 by 32% over the best of four baselines, VulLibGen, and its 110M-parameter KGC backbone alone already surpasses VulLibGen's best 7B-parameter configuration, with re-ranking adding consistent further gains across all evaluated LLM backbones.
Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment
Many cancer clinical trials fail from insufficient enrollment, and existing AI recruitment tools mostly assess eligibility in isolation without evaluation in real oncology workflows. TrialGPT 2.0 is a trial recommendation system that also judges which trials warrant consideration given a patient's current clinical needs and local workflow priorities, and it provides structured, inspectable explanations for expert review. In retrospective multicenter cohorts of 288 cases it surfaced at least one clinician-recommended trial in its top 10 for about 91% of cases while cutting clinician screening time by 55%, and a six-month prospective deployment in an active precision oncology tumor board expanded patient access to trial participation by 90.9% by catching opportunities the routine workflow missed. The authors also release NIH-TrialBench, 126 clinician-authored synthetic patient vignettes and matching scenarios from 11 NIH Institutes and Centers.
Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening
Crystal generators and tool-using agents now propose candidate structures faster than density functional theory (DFT) energy and phonon calculations can vet them, yet most screens check little beyond atomic overlap and give no chemical reason for rejection. Here agents generate, test, and actively refute two million candidate laws, distilling eight Plausibility Rules for Inorganic Structures (PRIS) that encode short-range repulsion, ionic contact and packing, electrostatic balance, bond-valence conservation, and crystallographic site complexity. Experimental structures satisfy the rule sets at 82 to 99% versus 6.5% for Pauling's rules 2 through 5 combined, and the strictest set detects 87.9% of damaged crystal structures where distance cutoffs catch only 1.6 to 3.2%. A derived synthesis score screens 83.7% of hard-to-synthesize structures while retaining 80.7% of experimental ones, cuts a DFT validation queue by up to 67.3% in an inverse-design run, and explains why GNoME is enriched in rare low-symmetry structures.
From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs
Scaling Transformers over user behavior sequences for production recommendation ranking faces two obstacles absent from language modeling: behavior signals are noisy, temporally irregular, and sparsely supervised, and each request must score many candidates against one shared user history under tight latency. ReST addresses signal quality with a sequence encoder using dual-gated attention, rotary positional and temporal embeddings, stabilized residual normalization, and training-only auxiliary objectives, and addresses compute asymmetry by splitting ranking into a heavy reusable encoder and a lightweight cross decoder with projection-free key-value attention, so a user prefix is encoded once and decoded many times. Across industrial and public benchmarks it scales more consistently along sequence length, depth, and width where LLM-style blocks saturate, and a one-week online A/B test on a production advertising platform raised online AUC by 1.31% and a core revenue metric by 11.93% within a 50 ms P99 budget, after which it was fully deployed.
Relational-Core Graph Analytics Querying graphs at SQL scale, and why the node/edge model is a performance tax, not a truer picture of connected data
A durable assumption holds that graph analytics needs a purpose-built graph engine and that relational systems handle connected data badly; the argument here is the reverse for the workloads enterprises actually run. ClickGraph and its Databricks-dialect sibling DeltaGraph translate Cypher directly onto existing relational schemas — the tables, columns, and foreign keys as they already stand — executing in place on ClickHouse, Databricks, or lakehouse files with no import and no separate cluster, which leaves ordinary SQL as an open optimization surface. The authors contend the node/edge property graph is a re-encoding of relationships relational tables already store explicitly, making query-time reconstruction pure overhead, and support this with a peer system's own published benchmark in which a columnar engine outruns Neo4j by two to four orders of magnitude plus reproducible runs across the LDBC Social Network Benchmark suite.
Can LLMs Design Video Coding Tools? A Case Study on Planar Mode
Designing video coding tools resists automation because any tool change couples tightly to the rest of the codec; the case study asks whether an LLM can redesign Planar mode, a long-standing intra prediction tool in video coding standards. The setup is a generation-and-evaluation loop in which the model proposes Planar predictors, encoder trials measure coding performance, and the model revises from that feedback. Replacing the default Planar mode in the Fraunhofer Versatile Video Encoder (VVenC) under its faster preset, the generated version achieves 0.18% bitrate savings for 0.4% complexity overhead on the standard benchmark. Extending to the Enhanced Compression Model (ECM), both replacing the newer directional Planar modes and adding the generated predictor as an extra mode with its own syntax elements produced gains in a constrained low-resolution setting.
StudentSim: Training LLM-based Student Simulators
AI tutors work best when they adapt to individual students, but evidence about which guidance suits which learner is slow and costly to gather, and existing student simulators either track state without processing explanations or role-play fluently without matching the target student's competence. StudentSim turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization, so each simulator both mirrors a student's own responses and updates them under tutor guidance. The accompanying StudentSimEval protocol covers 60 students across chess, second-language English writing, and mathematics and measures behavioral fidelity and guidance responsiveness; StudentSim outperforms GPT-5.4 on both metrics in all three domains, reaching fidelity 0.51 and responsiveness 0.91 in chess versus 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. Used as a reward model for tutor reinforcement learning, it produced a chess tutor that expert humans rated more accurate, better-guided, and more personalized than both a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward.
68 more specialized papers
- Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing L\'ea Bayati, Mohamed Dahmoune, Melek Rodoplu
- RAPIDMap: Rapid Multi-Agent Pipeline for Interpretable Disaster Mapping from Satellite and Street-view Imagery Yifan Yang, Lei Zou
- Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment Mustafa Talha \.Ilerisoy, Hung Manh Pham, Mathias Funk et al.
- ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation Yitong Han, Wei Gao, Yi Zhao et al.
- DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction Weiran Wang, Xintong Huo, Yueying Wang et al.
- Medical Causal Hypothesis Verification with Large Language Models Safiyyah Ahmed, Abrar Ansari, Md Aminul Islam et al.
- Life Operators: a self-evolving framework for multiscale life modelling Shuo Wang, Yike Guo
- MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts V. S. Anoop, Devika N
- AI Morbidity and Mortality: A Framework for Clinical AI Failure Review Paulius Mui, Dean F. Sittig, Steve Labkoff et al.
- Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models? Arkadiusz Lipiecki, Rafa{\l} Weron
- Generative artificial intelligence for reliable mechanistic reasoning for corrosion Bharath M N, R K Singh Raman, Alankar Alankar
- Intelligent Edge Computing Kalgi Gandhi, Minal Bhise
- Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking Md Rasel Khondokar, Qiao Qiao, Farjana Sultana Samia et al.
- Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer's Disease Detection Luqi Sun, Shreeram Suresh Chandra, Lin Zhang et al.
- Lightweight Adaptation of EEG Foundation Models for Stroke Motor Imagery Decoding: Domain Shift and Subject-Level Robustness Anh T. Nguyen, Zihua Sun, Michelle J. Johnson
- Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains Zi Wang, Minghui Xu, Tapan Mukerji
- TRUST: Threshold-Recalibrated Uncertainty-Safe Training for Certified Dismissal in Breast Cancer Screening Parham Hajishafiezahramini, Matthew Hamilton, Edward Kendall et al.
- A Human-AI Theorem Connecting Spontaneous and Field-Induced Mechanisms of Collective Behavior in One Dimension Weiguo Yin
- Latent-Space No-Arbitrage Geometry of Generative Models for Implied Volatility Surfaces Jing Wang, Shuaiqiang Liu, Cornelis Vuik
- From Tool Use to Technological Agency: LoopCAT as a Local-First, Open-Source Tool for Translation Technology Education Gokhan Dogru, Adri\`a Mart\'in Mor
- Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment Saad Mohammad Abrar, Eesha Kurella, Arnav Dadarya et al.
- Detoxifying Toxic Communication: A Design Science Approach to Responsible AI Hossein Arshadi Soufiani, Henry M. Kim, Hjalmar Turesson et al.
- NeuroPriv: Adversarial Representation Learning for Privacy in Wearable EEG Systems Sarmistha Sarna Gomasta, Bhawana Chhaglani, Prashant Shenoy
- A convolutional framework for detecting event-driven dynamics in energy price series Caixia Xu, Piotr Fryzlewicz
- A Multi-Branch Feature Fusion Approach for Health Misinformation Detection and Propagation Mkululi Sikosana, Sean Maudsley-Barton, Oluwaseun Ajao
- Accelerating Chemical Kinetics for Exoplanet Atmospheres using Neural Networks Isaac Malsky, Xi Zhang, Tiffany Kataria et al.
- Physiological Information Reliability: Cross-Layer Adaptive Resource Allocation for Cardiovascular Sensing Navaneeth Krishnan Kamalakannan, Janakiraman Kamalakannan, Harinisri Velmurugan
- AdaptNTK: Adaptive Uncertainty Quantification and Active Learning for Neural Network Potentials Prajwal Ananth, Shuwen Yue
- RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces Nikhil Wani
- VATO: A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows Xingxin Yang, Zhan Zhang, Yichen Li et al.
- CoVer: Conflict-Aware Claim Verification Shuning Zhang, Dai Shi, Bohao Chu et al.
- Learning Task-Specific Antibody Representations via Function-Aware Masking Ayan Goel, Thomas A. Walton, Amirali Aghazadeh
- DeSyR: A Decoupled Symbolic Recovery Framework with PINN-Guided Structure Search and Physics-Informed Coefficient Refinement Pancheng Niu, Jun Guo, Qiaolin He et al.
- GenONet: A Generative operator Network for High-Resolution Precipitation Nowcasting Mohammad Kian Golkar, Luciano Alves de Oliveira, Mohammad Khanjani
- BeamRMX: Radiation-Pattern-Driven Learning for Generalizable Beam Radio Map Prediction and Beam Management Yue Zhang, Xiucheng Wang, Wenshuo Chen et al.
- EEG-AS: Instance-Level Foundation Model Selection for EEG Foundation Models via Behavior Reconstruction Yunzhen Zhang, Ruoxi Piao, Hasan Onur Keles et al.
- HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields Lihao Chen, Xinyu Zhang, Panqi Chen et al.
- Visual Framing for News Stance Detection via Image Generation Dahyun Lee, Jiyoung Han, Kunwoo Park
- MaskCode: Mask Transformer for Feedback-Assisted Coding With Linear Block Codes Jonggyu Jang, Hongjae Nam, Vishrant Tripathi et al.
- Automated Tree Knowledge Graph Construction using Ontology Expansion and Retrieval from Vietnamese History Textbooks Ket Doan Nguyen, Minh N. H. Nguyen
- Ctrl-F-Resist. Practices, Challenges, and Technical Needs of Civil Society Organizations Monitoring the Far-Right Online Elisabeth Steffen, Helena Mihaljevi\'c
- TWIX: a Two-Stage Approach for End-To-End Named Entity Recognition and Relation Extraction Marco Martinelli, Laura Menotti
- Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources Ivan Decostanzi, Michele Ronco, Sergio Consoli et al.
- FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study Sai Puppala, Koushik Sinha
- iPINN for Broadband CARS Phase Retrieval: A Framework for Function Approximation and Inverse Modeling Problems in Nonlinear Spectroscopy Ravi Teja Vulchi, Carl Messerschmidt, Mohammadsadegh Vafaeinezhad et al.
- Direct Optimization of a 3D Finite-Source Reflector via Neural-Network Parameterization Roel Hacking, Lisa Kusch, Martijn Anthonissen et al.
- PersianAnonymizer: Evaluating LLM-Labeled Training for Efficient NER-based Anonymization in Persian Mohammad Hossein Shalchian, Mostafa Amiri, Amir Mahdi Sadeghzadeh
- Few-Shot Out of Domain Intent Detection with Covariance Corrected Mahalanobis Distance Jayasimha Talur, Oleg Smirnov, Paul Missault
- The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research Yuni Susanti, Moritz Schubotz
- On the Human and Computer Alignment of Attribute-Based Music Matches Roser Batlle-Roca, Woosung Choi, Joan Serr\`a et al.
- A Network Science Perspective on Evaluating Deep Graph Generative Models Tianrui Mao, Abele Malan, Megha Khosla et al.
- Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech Che Hyun Lee, Sangkwon Park, Donghun Kang et al.
- Web Price Extraction: State of the Art and an Adaptive Browserless Implementation Evgeniia Kositsyna, Jorge Lloret-Gazo
- User Representation via Cross Multi-source Behavior Pre-training for Mobile Games Chengqi Yang, Yiran Qiao, Feng Liu et al.
- Space Generative AI with Solar Energy Harvesting Jierui Zhang, Jianhao Huang, Zhanwei Wang et al.
- Text-guided flow matching enables sample-efficient crystal structure generation Wentao Li
- PersuaRL: Reinforcement Learning-Driven Multi-Expert Selection for Persuasive Dialogue Generation in Insurance Rohan Kirti, Akash Ghosh, Aryan Vats et al.
- Analog-DB: An Agent-First Analog Integrated Circuit Database, From Blocks to Systems Danial Noori Zadeh, Mohamed B. Elamien
- GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri et al.
- Automated Event Log Generation from Unstructured Text Using Finetuned LLMs Maximilian Seeth, Gabriel Marques Tavares, Daniel Schuster
- SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding Handong Wang, Jiaxin Qi, Baisheng Lai et al.
- PopPert: Population-level Joint-Distribution Modeling for Single-Cell Perturbation Prediction Handong Wang, Jiaxin Qi, Haochen Feng et al.
- Predicting Subsurface Abnormalities Growth using Physics-Informed Neural Networks Mehrdad Shafiei Dizaji, Hoda Azari
- CATeye: Coupled Attribute-Topology Invariance Learning for Voucher Abuse Detection Tian Tian, Shuaicheng Niu, Hao Kuang et al.
- Semantic-Guided Multimodal Preprocessing for Vision Transformer-Based Clear Cell Renal Cell Carcinoma Grading Fatemeh Javadian, Zhu Chen, Zahra Aminparast et al.
- Quantum Sparse Autoencoders for Q-Matrix Estimation in Cognitive Diagnosis Arif Hassan Zidan, Yi Pan, Bowen Guo et al.
- A systematic Approach to constructing a Chance-and-Risk Matrix for Semiconductor Supply Chains Ema Salki\'c, Alexander Fichtl, Philipp Ulrich et al.
- Designing Proactive Thought Partners for Writing Chao Zhang, Abe Davis, Chih-Wei Chen et al.
Agents 82
HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models
Text-based world models must learn symbolic action effects from serialized state descriptions, but how that state is formatted has been largely unexamined. HyperWorld compares raw observations against three symbolic serializations of identical ground-truth state — independent sentences, pairwise triples, and entity-centered hyperedge units that bundle several related facts around an entity or relation — under one training objective that predicts effects or flags an action infeasible. Hyperedge grouping helps most at 0.5B–1.5B parameters and under distribution shift, taking the best out-of-distribution fact F1 and the highest success rate in downstream greedy planning, while larger models narrow the gap and pairwise triples can edge ahead on in-distribution exact match.
Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
Long-horizon agent benchmarks report end-to-end success but mix state-tracking difficulty with instruction ambiguity and can be gamed by a hallucinated final answer, leaving it unclear whether a model can carry exact intermediate state at all. The test here is computing an MD5 hash step by step: 196 dependent tool calls over 64 rounds, with the model holding four 32-bit words in its own context between calls, checked against a from-scratch RFC 1321 reference trace so any error is pure bookkeeping. gpt-oss-120b, a mixture-of-experts model with only about 5.5B active parameters per token, carries state across all 196 calls and returns the correct digest on a majority of completed runs, including a variant where every arithmetic primitive is replaced by a second LLM worker; success hinges on keeping the model's own reasoning in context each turn and voting over a thinking-enabled worker to cancel modular-arithmetic slips.
OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets
As deployments grow into fleets where several agents, planners, and execution backends touch the same environment, safety becomes a question of whether a concrete action should be allowed to commit rather than whether a prompt looked risky. OpenAgentFlow splits the system into a control plane and an action plane, normalizing pending GUI actions, API calls, tool calls, and LLM-generated invocations into one AgentEvent stream that passes through a shared pre-execution policy enforcement point, with provenance, session state, audit records, and hot-updatable policies held centrally so new rules apply without touching agents, prompts, or models. An Android instantiation reaches 94.0% accuracy and a 95.3% attack block rate on a 300-case action-event benchmark, matches expected behavior in 27 of 30 dynamic-policy cases after new rules are installed, and scores a 92.9% trace-adjusted pass rate on emulator traces spanning GUI, API, and planned actions.
UI-Venus-2 Technical Report
Graphical user interface agents tend to be tuned for benchmarks rather than deployment, held back by narrow environment coverage, brittle task construction, and reward signals that cannot be trusted. UI-Venus-2 scales three axes together for a single closed-loop reasoning-and-action foundation agent spanning mobile, web, and desktop: coverage of more than 170 multilingual mobile apps plus native desktop operating systems, a deep-research pipeline that generates function-grounded instructions, and trace-level plus sample-level verifiers using visual keypoints and multi-model voting to produce reliable reinforcement learning rewards. Safety-aware mechanisms gate consequential actions, and the model is released open source.
EULER: Exploring Underused Links with Evidence-Checked Return for Multi-Agent Mathematical Discovery
Moving a mathematical problem between communities that use different objects, invariants, and tools is costly, so such transfers are usually skipped. EULER makes that transfer — a bridge — the unit of search in a multi-agent system: direct, adjacent-domain, and distant-domain routes compete for budget around a fixed conjecture, and a bridge keeps its budget only if it supplies an operation the source representation cannot execute and its target-side evidence returns to the original statement through a checked implication, with six ordered stress tests filtering invalid bridges before expensive search. On 120 recent combinatorics conjectures frozen and screened for contamination, the system produced 10 proofs, 3 refutations, and 45 scoped partial results; ablations show the stress tests cut incorrect conclusions from 9 to 3, and success tracked executable operation gain and valid return rather than domain distance.
trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
Production evaluation of LLM agents usually shows a judge only the request and the final reply, which is structurally incapable of noticing an agent that got the right answer the wrong way. The setup measures that blind spot with ground truth by construction — a deterministic tool-using support-desk environment, a scripted oracle policy, and a fault injector that breaks exactly one thing at a known step, with faults stratified by whether the customer-visible outcome survived. Across 400 trajectories and five judges, the outcome-only judge catches 84% of outcome-breaking faults but only 45% of silent ones while falsely flagging 33% of correct trajectories, whereas a step-rubric judge reaches 77% silent recall with zero false alarms at three times the cost; no judge reads the final reply, so an invented promise appended to a perfect trajectory slips past the rules entirely and the step judge 82% of the time, and self-consistency triples cost without improving anything.
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
Models that generate the next graphical user interface screen are usually scored one step at a time, even though their purpose is to serve as multi-step environments where generated states get reused as the basis for further interaction. GUI-CC tests that reuse directly with two tracks: an offline track rolling models along 500 real mobile trajectories drawn from GUIOdyssey, and an online track where fixed probing agents interact with model-generated interfaces on 200 emulator-verified tasks across 30 apps, scoring transition fidelity, transition plausibility, contextual consistency, and task progress. Plausible single-step generation turns out not to imply reliable simulation — current models render usable-looking screens while losing task-relevant context and failing to support executable multi-step rollouts.
Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness
Autonomous agents that modify infrastructure, deploy services, and verify their own results need explicit machinery for progressing through long tasks, containing what they can execute, and recovering from failure. The proposed framework separates three concerns: graph engineering encodes workflow progression with verification-gated transitions, loop engineering bounds diagnosis, repair or re-planning, retry, and re-verification, and an agent harness enforces zero-trust execution through identity, authorization, policy-scoped capabilities, isolation, and runtime safeguards. Instantiated on Google Cloud, every run terminates either in a verified operational deployment or an auditable terminal failure within its recovery bounds, with progression requiring machine-checkable repository, deployment, and runtime evidence.
AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes
Commercial large language model (LLM) APIs may silently substitute, quantize, or wrap the backbone they advertise, and every existing audit infers identity from the text channel — which agentic serving stacks discard once the model emits a tool call, and which provider-injected system prompts can distort enough to falsely accuse honest providers. AgentProv instead fingerprints a deployed model by its categorical tool-call distribution and decides identity with a maximum mean discrepancy (MMD) permutation test, on the observation that agentic post-training writes tool-use behaviour into the weights in a way that survives deployment context. It caught every substituted model across 630 evaluated checkpoint pairs, a 100% detection rate, while holding the false-positive rate under system-prompt injection to 7% versus 67% for MET and 53% for RUT; on third-party endpoints its disagreements with MET line up with a token-count side channel that reveals injected system prompts.
CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
Writing fast CUDA kernels requires algorithm design, correctness validation, and hardware-aware tuning, and prior LLM work mostly transpiles from PyTorch to CUDA rather than generating kernels directly from natural language (Text2CUDA), where the model must bridge high-level intent and low-level implementation. CUDA-Harness introduces Intermediate-Structured Generation to connect semantic understanding with kernel code, Synthesis-Based Verification that supplies isolated test data and progressive validation to blunt the reward hacking that comes from relying on predefined test inputs, and Feedback-Adaptive Evolution that prioritizes correctness before optimizing speed. Experiments report gains over prior approaches along with generalization across different LLMs, hardware platforms, and C-to-CUDA transpilation.
A Formal Analysis of Agent Payment Protocols
Agent payment protocols let AI agents buy goods and settle payments for users, spreading intent, delegated authority, credential use, settlement, and fulfilment across actors and stages in ways that no single message can enforce, yet their security guarantees remain implicit across specs and reference implementations. Four representative protocols — x402, MPP, ACP, and AP2 — are modelled in the Tamarin prover under a shared abstraction of the payment lifecycle, using source-backed verification questions and counterexample traces rather than a presupposed property taxonomy, yielding 18 shared security principles. Across 86 verification cases the analysis reproduces 46 known or calibration cases and surfaces 40 previously undocumented formal-consistency findings, each traced to a missing protocol relation and reverified against a minimally strengthened model, with ten validated through proof-of-concept exploits, schema-level witnesses, and executable traces.
Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents
An agent asked to analyse an experiment will usually produce running code, but whether the analysis is defensible depends on procedural choices — which statistical test the field accepts, which identifier namespace is authoritative, which caveats must accompany a result. Scientific Agent Skills is an openly licensed library of 163 such procedures across 16 areas of practice, spanning genomics, cheminformatics, medical imaging, study design, and scientific communication. Each skill is a directory around a versioned human-readable instruction file that the agent loads only when a task calls for it, often alongside reference material and runnable scripts; the authors report no task-level evaluation and no measurement of how often hosts select a skill.
AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis
Powder X-ray diffraction (XRD) analysis resists automation because an agent must read diffraction evidence, drive refinement software, order coupled parameters defensibly, and tell a numerical improvement from a physically valid one. AutoXRD structures the task as stepwise refinement grounded in observed evidence with deterministic crystallographic and physical checks gating acceptance, and XRDBench evaluates it on 100 bounded reasoning tasks plus 34 executable end-to-end workflows requiring file inspection, software execution, iterative refinement, and reporting. Across 1,340 model-task runs, ten recent LLMs average only 57.8 out of 100, dropping from 61.9 on the reasoning track to 53.7 end-to-end, with GPT-5.6 Sol highest overall at 81.1; execution traces expose recurring failures in coupled-parameter control, quantitative reasoning, evidence preservation, and knowing when to stop.
Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops
Code-level autonomous research loops, where a large language model (LLM) agent proposes edits to a training pipeline, runs it, and keeps changes that improve a measurable in-loop metric, are examined here for whether those gains reflect genuine progress. Across several experimental settings the authors identify a failure they call algorithmic mode collapse: edits keep touching different lines of code while the underlying kinds of algorithmic change repeat, and in-loop gains diverge increasingly from held-out evaluations. Their mitigation, DAPS (Diversity-Aware Proposal Sampling), combines category-coverage reweighting, a persistent edit memory, and a validation gate, and under a three-tier protocol separating in-loop, audit, and blind metrics it cuts semantic-cluster decay of edits by 69.1% and improves relative faithfulness by 83.7% on the blind metric while preserving in-loop optimization speed.
Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations
FAIRY is a full-stack agent system deployed on an operating soybean research farm at Harbin Institute of Technology, covering a whole season from ridge preparation and planting through irrigation, fertilization, pest and disease treatment, harvest, grain handling, drying, and storage. It integrates production machinery APIs, fixed soil and canopy sensors, multispectral and thermal drones, satellite vegetation products, a weather station, calibrated crop-process models, and multi-season yield histories under an "everything is an event" execution paradigm, on top of which sit a library of atomic agronomic skills, multi-agent controllers, frontier and edge model execution, and full-path trace logging. The authors evaluate nine state-of-the-art agent controllers across one hundred full-season scenarios on a 64-ridge field, scoring agentic success, spatiotemporal correctness of the entire action path, token cost, and edge-device runtime.
ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation
Slide generation from documents demands both faithful content selection and precise spatial layout, but current slide agents rewrite a whole slide or deck, render it, and only then critique — delayed feedback that makes local failures like overflow, overlap, clipping, and off-canvas placement hard to attribute or repair. ReDeck breaks revision into atomic edit actions and returns renderer-derived observations after each one, turning the loop into "one edit, one observation," layered with turn-level adaptive critique for semantic and design guidance and a submission-level gate enforcing hard layout validation. Paired with DeckQuiz, a benchmark separating content fidelity, spatial correctness, and design quality, ReDeck outperforms existing slide agents across GPT-5.4, Claude-4.6, and Gemini-3.1, with ablations showing feedback timing and granularity both matter.
WHALE: A Simple Recipe for Joint Harness-Weight Optimization
An agent's performance depends on both its model weights and the executable harness that manages context and control flow, and tuning either alone leaves the system bottlenecked by the frozen half — weight updates change which harness works, and harness changes expose different model capabilities. WHALE (Weight-Harness Alternating LEarning) alternates two phases, updating weights under the current harness via online rejection-sampling fine-tuning, then searching for a better harness under the updated model via Meta-Harness, switching on either fixed phase durations or an adaptive patience rule. With Qwen3.5-2B/4B agents on search question answering, mathematical reasoning, and chess puzzles, it beats weight-only, harness-only, and Fast-Slow Training by 4.15 to 24.38 percentage points in best mean@8 accuracy, and the bottleneck genuinely varies: harness search matches peak weight-only accuracy with far fewer rollouts on SearchQA, while math improves only after a weight update.
Don't Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field Changes
When a language model agent proposes a fix in a GitOps workflow, applying it means editing a version-controlled config file, and the natural implementation — having the model write the diff or the new file — turns out to be unsafe for unattended automation on real Kubernetes manifests. Under strict patching almost no unified diffs apply, while a tolerant tool like GNU patch applies 96% but silently misapplies roughly 1 in 7 (14-20%) with no error signal; full-file rewrite is capability-dependent, with a small model corrupting files and a frontier model usually correct but nondeterministic and costing O(file size) per edit. The proposed alternative has the agent emit only a structured field-change intent — which resource, field, and value — while a deterministic pipeline indexes manifests by kind and name, locates the target scalar's exact character span through the YAML parser's node position marks, and replaces just that span in the raw text, preserving comments and formatting at O(1) generation cost. It ships as KubeAstra under Apache-2.0 with the benchmark released, scoped to faithfully applying a known change rather than judging whether the change is correct.
Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems
Orchestrating multi-agent language model systems requires deciding what each agent does next, and both obvious options fail: routing on the query alone cannot react to intermediate progress or errors, while routing on the full execution history forces every later decision to reprocess redundant steps, inflating cost. Gated-Memory Routing conditions each decision on the query plus a learned execution memory, where a Memory Write Gate commits only non-redundant reasoning steps, a Retrieval Gate hands each agent a compact relevant subset, and an Adaptive Halting Controller stops once the memory holds enough evidence to answer. Across five reasoning and code-generation benchmarks it achieves the best average accuracy, beating the strongest baseline by 2.44 points while cutting HumanEval inference cost by 31.9%.
Invalidation Contracts for Cross-Episode Agent Memory
Agents that cache recovery advice from API errors save tokens across episodes, but server-side data changes silently turn those cached fixes wrong, and re-deriving every time erases the savings. Invalidation contracts attach version stamps and cacheability hints to each recovery suggestion so a client can evict exactly the stale entries, and they split realized savings into validity (fraction of cached fixes still correct after drift, a property of the protocol alone) and compliance (fraction the planner actually applies first try, a property of the model). Across seven models, three serving paths, and roughly 9,400 episodes, row-level invalidation raises compliance by up to 66.7 percentage points and recovers 29-33% of baseline token cost on four models, while table-level invalidation drops post-drift first-try rates to zero on five of seven. Identical wire bytes produced 100% first-try compliance on Claude Haiku 4.5 but 11% or below on Claude Sonnet 5, which refused fixes adding fields absent from the original request.
Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems
When agents hold credentials, call services, and spawn sub-agents, the distributed-systems question of who may act on whose authority becomes acute because the deciding component is a language model an adversary can hijack. The authors argue for evaluating agent security under an untrusted-model assumption, where a fully prompt-injected agent still cannot exceed its explicitly delegated authority, and derive eight requirements from four adversaries: confused deputy, token theft and replay, prompt-injection privilege escalation, and compromised sub-agents. A default runtime using broad bearer credentials with authorization decided inside the model fails all four, and among LangGraph, CrewAI, AutoGen, and the Model Context Protocol authorization model, three offer no built-in confinement and one only partial. Their authorization broker blocks all four threats, accepted 0 of 200,000 forged tokens, confined a compromised sub-agent to a mean of 1.5 reachable actions versus all 8,100 under bearer delegation, and costs about 2.6 microseconds per decision.
The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems
Agent fleets now take actions that cannot be undone — moving money, deploying code, deleting data, disclosing information — and because current controls authorize one effect at a time, a set of individually correct decisions can still exhaust a principal's overall risk tolerance under a shared trigger. The proposed irreversibility budget makes residual value-at-risk a first-class resource that a trusted runtime accounts for per principal across agents, workflows, and tenants, charging each effect its residual loss and denying the marginal action once the aggregate would overdraw. In a controlled study, per-effect gates permitted fleet-level overdraws of up to 48 times the tenant's risk limit while the budget kept every correctly priced run inside it. The authors name conservative, dependency-aware pricing of heterogeneous, adversarially declared, correlated effects as the open problem blocking deployment.
Toward Workflow-Aware Benchmarking for Healthcare NLP Agents
Healthcare agent evaluations mostly use static medical question answering or one-shot generation, which omit longitudinal state, interruptions, and handoffs to humans. The proposed episode-level protocol separates evidence attributable to the model, the agent, and the simulated workflow; specifies a five-field episode schema; and defines annotation and scoring for state continuity, evidence traceability, and escalation decisions, with cost-sensitive treatment of missed versus unnecessary escalation. It is instantiated as four task templates — documentation update, evidence retrieval, patient messaging, and triage handoff — and is explicitly framed as a reproducible intermediate layer between static benchmarks and prospective workflow studies, not a measure of clinical outcomes.
Dr. Claw: An AI Scientist Workspace for Vibe Research
Command-line coding agents can already read and write files over long sessions, yet research work still scatters across chat tools, IDEs, terminals, and writing environments, and the decisions that would make it auditable go unrecorded. Dr. Claw is an open-source workspace that wraps existing coding-agent executors in a human-in-the-loop workflow rather than adding another autonomous agent, using persistent state objects, a reusable skill library, and multi-executor coordination to tie planning, execution, and writing into one traceable and recoverable loop. Evaluated against a bare command-line agent sharing the same backend executor, so the comparison isolates the orchestration layer, the wrapped system scores higher on research completeness while leaving an auditable, recoverable process trail; the paper also walks through an interactive three-view scenario and a failure-recovery case.
FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
Origami knowledge circulates mostly as unstructured demonstration videos, while computational tools need structured representations such as crease patterns or executable parametric plans. FoldingAgent bridges the two with a vision-language model agent that infers explicit parametric folding programs from video, using specialized tools to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions over a parametric space of paper geometry plus folding actions. Because the agent acts sequentially and can re-plan, it mitigates the error compounding that plagues multi-step folding. On PurelandFold, a newly curated benchmark of Pureland origami videos with ground-truth geometry and action labels, the combination of model reasoning, tools, and physical simulation converts unstructured demonstrations into executable, physically plausible folding procedures.
RestoreBench: Can AI Agents Restore Power Flow Convergence?
Working out why a power flow case fails to converge and fixing it takes engineering judgment, experimentation, and iterative decisions inside a constrained action space — a plausible but largely untested target for tool-using LLM agents. The benchmark specifies the simulation environment, observation and action spaces, and evaluation metrics over two power grids with 46 non-convergent cases each, every case requiring one or more corrective actions to restore convergence. Several large language models are compared across three architectures — plain chatbot, single agent, and multi-agent — giving a reproducible starting point for agentic systems in power system planning and operation, with code released publicly.
SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation
Spectrum management decisions increasingly require pulling together policy proceedings, legal regulations, and license databases that are disaggregated, mix text with tables, and are formatted for human readers rather than machines. SpecMind is a multi-agent retrieval-augmented generation (RAG) system where coordinating agents dispatch specialized sub-agents to retrieve and synthesize across these heterogeneous sources, evaluated on SpecBench, a new question-and-answer dataset built from real license records and policy proceedings to fill a gap in domain evaluation resources. The system reports over 80% win rate against strong general-purpose RAG baselines across spectrum tasks, which the authors attribute to more accurate retrieval and better contextual reasoning from the agent-based decomposition.
SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents
Judging a task-oriented dialogue turn means checking whether it actually advanced the underlying workflow state, something holistic LLM judges can miss because they weigh the whole context at once and need at least one full model call per turn. SAGE compiles a workflow specification and per-turn state diff into atomic, schema-grounded criteria, then routes each through a cascade of symbolic rules and on-device encoder/natural-language-inference verifiers that abstain rather than guess, aggregating criterion verdicts into a turn-level decision with an evidence trace. Its recommended SAGE-Core operating point settles 81-91% of criteria at zero paid model cost, and across four slices of MultiWOZ, Schema-Guided Dialogue, and ABCD no LLM-as-a-judge baseline significantly beats it — including a state-aware GPT-4.1 judge costing $4.7-8.0 per 1,000 turns. A two-annotator audit (n=200, kappa 0.94) supports label fidelity on transcript-visible failure classes, while the authors scope construct-validity limits from injected failures and partial symbolic circularity.
mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers
Handing an agent a dossier about a named expert is often claimed to do three separate things at once: supply hard-to-find material, produce a recognizable persona, and improve the agent's judgment. mimeo is an open-source tool that gathers a person's public work, verifies each extracted quotation against cached source text (rejecting 13.2% of them), and emits a loadable agent file, tested here across four expert profiles in one coding-agent harness. Knowledge access was the clearest win, with mimeo answering all 20 obscure quotation-heavy questions while no closed-book condition exceeded 10, though BM25 keyword search over the same pages answered 15-17; grounding also prevented the factual misstatements that memory-written personas made on 1-4 of 20 answers. Judgment transfer stayed unresolved because every condition hit the ceiling on the engineering and application tasks, and an AI judge scoring whether answers sounded like the expert disagreed across judges, which the authors flag as a warning against single-judge evaluation.
Towards a Belief-Based World Model for LLM Agents
Large language models used as decision-making policies struggle on long-horizon tasks under partial observability, and the usual world-model remedy, simulating candidate actions before committing, says nothing about how uncertain the agent is about the current state. Belief-Based World Models (BB-WMs) instead maintain an explicit belief the policy can query to see what is known and what is unknown right now. Before tackling how to learn such beliefs accurately, the authors test the prerequisite question and find that exposing a world model's belief directly to an LLM policy improves task performance under partial observability, with gains that remain complementary to simulation-based world models.
Exploring Collaboration between a language and a non-language agent
When a language model orchestrates specialist subagents, any non-language expert such as a chess engine or a robot controller must have its rich continuous state compressed into a short text summary at every step, and it is unclear how much that costs. LLAMIA-Bench measures it with six collaborative chess tasks covering behavioral imitation, state assessment, and natural-language explanation, each a problem neither the LLM nor the engine solves alone. The alternative proposed, latent state internalization, projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens that are re-encoded as the position changes; against verbalized integration this reveals a verbalization debt that widens over training and persists from 4B to 14B parameters. The resulting 14B LLAMIA model matches or beats task specialists and frontier systems including GPT-5.1 with tool access on every task, and holds up out of distribution where task-specific finetunes collapse.
Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications
Computer-use agents that combine language reasoning with visual interface grounding are pitched as a general way to operate desktop graphical user interfaces, but whether they help blind screen-reader users in real workflows had not been measured. A three-week diary study put OLLA, a screen-reader-accessible agent prototype, in the hands of 8 blind participants, capturing 1,258 commands across 12 applications together with screenshots, accessibility trees, model responses, and action traces, then replayed the same commands through four additional models. GPT-5 led with a 52.5% success rate, and trace analysis attributes the remaining failures to grounding, planning, constraint tracking, and knowing when to stop, while interviews surface user needs that full automation does not address.
Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers
Agents are usually described by whichever model and harness currently runs them, which leaves no vocabulary for an agent that must survive model swaps, orchestration changes, and host migrations while keeping one identity. The proposed architecture separates a continuity-bearing substrate (identity representation, durable private memory, versioned software body) from a replaceable deployment binding (reasoner, harness, host, and interaction surfaces), defines six continuity invariants, and specifies a quiesce-checkpoint-validate-bind-rehydrate-resume migration protocol. A reference implementation called Enoch passes 833 core tests plus 92 provider and library tests in a clean-room run and has survived reasoner-version, interaction-surface, and host-machine substitutions, which the authors are careful to frame as evidence of mechanical substitutability rather than behavioral invariance.
Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents
Agents that retrieve external skills are typically scored by comparing tasks where retrieval fired against tasks where it did not, a comparison contaminated by selection bias. The authors define Skill Following and measure it with the Retrieval-Invoked Actual-Use Effect (RAE), which compares matched skill-enabled and skill-disabled runs of the same task, restricted to tasks where the agent actually retrieved a skill. Across 17 large language models on coding and math, many models show a positive aggregate retrieval lift while their RAE is negative — on MBPP+, several models that look better system-wide are in fact hurt on exactly the tasks where retrieval happened.
WiseSpec: Requirements-Driven Agents for Code Generation
Coding agents often fail on repository-level tasks not because their tools are weak but because the task description itself is incomplete, ambiguous, or missing context. Borrowing from software requirements engineering, WiseSpec automatically builds structured, information-rich requirements before code generation, scores their quality through execution-based evaluation, and iteratively refines them to better steer the generator. The framework beats all baselines tested, with an average improvement of 13.17% in percentage of issues resolved.
Investigating Assistant Bias in LLM User Simulators Using a Role Vector
LLM-based user simulators used to evaluate autonomous agents suffer from "assistant bias": they stay cooperative and goal-directed instead of reproducing the frustration and disengagement of real users, which undermines evaluation validity. The authors extract a user role vector from model activations by contrasting how the model encodes the user versus the assistant perspective on the same dialogue. They find the user direction is linearly identifiable, elicits user-like behavior when steered, and captures traits distinct from assistant ones, but that amplifying it exaggerates user behavior and can override the specific user profile being simulated.
Towards Effective Structured Context Modeling for Conversational Recommender Systems via Dual-node Monte Carlo Tree Search
Conversational recommender systems must both elicit user preferences over multiple turns and exploit them for retrieval, but typically model dialogue context as flat history. DREAMS structures context as a tree with two node types: elicitation nodes that use Monte Carlo Tree Search (MCTS) to explore which conversational actions best reveal latent preferences, and exploitation nodes that use LLM refinement to convert the tracked preference state into structured retrieval queries. Experiments on benchmark datasets support the effectiveness of the dual-node design.
Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
In multi-agent LLM pipelines a single prompt usually carries two entangled jobs — producing task content and specifying execution protocol such as message routing, output format, and termination signals — so an optimizer tuning the content can silently break the protocol and crash the pipeline. The proposed control-data flow separation encodes execution-critical control as typed, validated program objects while leaving only natural-language task content exposed to prompt optimization. Across synthetic reasoning, collaborative review generation, and insurance rating workflows, the framework achieves 100% eventual protocol validity while still improving task performance.
REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows
When a user revises their request mid-run, an agent workflow must choose between discarding in-flight work (correct but wasteful) and reusing it (fast but risking stale state leaking into outputs and tool side effects). REVISE is a runtime that intersects the revision's delta with recorded data and control dependencies, propagates the impact through the partially executed DAG, halts only the invalidated work, and recomputes just the affected region while revalidating reused results before commit. Analysis of real coding-agent traces shows substantial overlap worth salvaging (56.55 s enqueue-to-completion overlap at p95), and across 300 revision executions it matched a latest-version oracle with no stale outputs while cutting model calls by 40.6–56.0% versus full restart on unmodified LangGraph and LLMCompiler applications running Qwen3-14B.
Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search
Language model agents that propose actions, observe feedback, and explain themselves offer stated confidence and rationales as cheap monitoring signals, and the authors test whether those signals hold up against ground truth. A model drives an evolutionary search over Contexto, a word game whose feedback assigns every valid guess an exact rank without human annotation, producing 12,249 self-reports across 200 runs, five configurations, and three model families. All three tested assumptions fail: operators overstate their top-100 success rate by factors of 4.8 to 9.3, controlled swaps of 754 inherited rationales bound any genuine benefit at roughly 250 ranks, and fitness-based selection shows no detectable improvement in report accuracy over random selection despite producing very different search behavior.
Creative Generation via Multi-Agent Debate: Does Debate Suppress Diversity?
Multi-Agent Debate improves answers on factual and reasoning tasks by pushing agents toward agreement, but that same convergence collapses variety across independent runs, which matters for narrative writing and scientific ideation where exploration is the point. The authors show that keeping agents divergent within a single debate session is a necessary condition for diverse outputs across runs, then build Creative-MAD on two mechanisms: Cognitive Lens Assignment anchors each agent to a distinct persistent cognitive mode to counter identity drift, and Embedding-based Peer Selection limits each agent's context to its most semantically distant peers to counter majority pull. On four creative benchmarks this raises both lexical and semantic diversity while preserving debate's quality gains.
ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything
Building multi-agent systems on top of language models forces a choice between code frameworks that are expressive but engineering-heavy and no-code builders that lock agent interactions into workflows the author has to define up front. DevAll pairs a declarative executable graph abstraction with a cycle-aware execution engine, so heterogeneous agents and dynamic, cyclic interactions fit into one representation, and wraps it in a visual interface for authoring, running, monitoring, and inspecting systems — human-in-the-loop steps included — entirely without code. Experiments show it reproduces state-of-the-art multi-agent systems on three representative tasks at competitive performance with no task-specific orchestration code, and the platform ships as part of the open-source ChatDev repository.
ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents
Long-horizon language-model agents must decide what goes into each prompt, in what order, and when to compact history under a hard context window and a byte-sensitive prompt cache — logic that in production is scattered across prompt builders, ad hoc compaction, cache-break workarounds, and per-provider shims. The authors argue this is structurally the same problem as relational query execution, and build ContextPipe around that analogy: a five-phase Plan-Bind-Optimize-Execute-Feedback pipeline over a structured data-source catalog, with a deterministic cache-aware optimizer and an EXPLAIN ANALYZE-style trace that makes context auditable, replayable, and failure-isolated. A preliminary run on the Qutebrowser subset of SWE-bench Pro cuts total token volume by 31% against append-only context construction, along with 23% fewer model calls and 9% lower response time, at the cost of a worse KV cache hit ratio.
Agentic programs: an emerging form of scientific software in computational materials science
Computational materials science has conventionally handed algorithmic steps to computers and kept scientific judgement with humans, a split that current agent harnesses make negotiable. The authors argue for a category they call agentic programs: scientific software that couples deterministic algorithms with bounded LLM-based judgement, task-specific verification, episodic maturation of the harness, and full delegation once in production. The concept is illustrated with DeMARS, an agentic program that constructs atomistic models from experimentally measured disordered crystal structures.
One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning
Search agents trained with reinforcement learning are normally optimised against a single fixed tool-call budget, so they degrade when the deployment budget differs from the training one. AnySearch trains one policy to handle any budget in two phases: first with a scaffold that injects the remaining budget into the agent's state and prompts structured reasoning about allocation under linearly decaying budgets, then with the scaffold removed and budgets sampled adaptively to match inference conditions. A composite reward couples answer accuracy with budget efficiency, weighted so the efficiency term is amplified on queries the agent answers well and damped on ones it does not. Across seven general and multi-hop question-answering benchmarks the single policy beats baselines at every budget scale and generalises to constraints outside its training range without inflating token usage.
Polished but Unresolved: Identifying Late-Stage Pressure States in Long-Horizon Tool-Use Agents
Long-horizon tool-use agents fail not only by searching or planning badly but by stopping too soon, submitting an answer that looks finished while constraints remain unmet. A linear probe trained on the agent's hidden states shows this "late-stage pressure" state is linearly identifiable from activations, and steering the hidden states along the identified direction changes both the probe's pressure score and whether the agent keeps calling tools or submits. Controlled context manipulations further show the pressure eases when constraints are stated clearly and actions are mapped explicitly. Those findings motivate Probe-Sensed Pressure Relief (PSPR), a plugin that applies a light steering nudge under moderate pressure and switches to structured organisation under high pressure, improving several existing agent methods on long-horizon benchmarks.
HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution
Agents that rewrite their own harness — prompts, skills, tools, execution logic — from environment feedback run into three failure modes: terminal pass/fail signals make it unclear which step caused an error, agents memorise task-specific tricks instead of general capability, and unguarded edits erase earlier competence. HarnessEvolve separates execution from evolution across independent modules for execution, evaluation, optimization, and gating, and attacks credit assignment by generating reference trajectories (runs performed with the ground-truth answer in hand) and aligning failed runs against them to extract and cluster systematic error patterns. Every candidate harness edit must clear both a quality gate that filters data leakage and prompt bloat and a performance gate that requires improvement on the current batch without regression on recent ones, with held-out validation at epoch end picking the best accepted snapshot. Across open-domain and enterprise benchmarks, models, and agent frameworks it reports consistent gains over state-of-the-art self-evolution baselines.
Dense Process Supervision for Search Agents via Fact Utility Estimation
Reinforcement learning for search agents typically rewards only the final answer, leaving no signal for which intermediate retrieval steps actually helped. The proposed method models reasoning as the accumulation of discrete evidence facts: structured facts are extracted from raw observations into an explicit fact store, semantically equivalent facts are clustered, and Bayesian estimation over group rollouts infers each cluster's posterior utility, which is converted into dense step-level rewards for training. Across seven single-hop and multi-hop question-answering benchmarks the approach consistently beats outcome-reward baselines, with ablations showing the clearest relative gains on multi-hop questions where credit assignment is hardest.
Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems
Optimization solvers handle vehicle routing problems well, but translating a complex real-world variant into solver input demands expertise that keeps the technology out of most hands. RLEA automates that modeling step with a multi-agent framework in which a lightweight neural planner, trained with soft Q-learning, orchestrates the actions of language-model agents, backed by an evolutionary memory module and retrieval over external solver documentation for program generation and refinement. Evaluated on 48 distinct routing variants across several solvers, it reaches a 16.67% higher success rate than the previous state of the art while producing substantially fewer runtime errors.
MemoryWalker: Stop Training Agents on Contexts They Never Saw
Agent harnesses such as Claude Code and Qwen-Agent compress context mid-rollout, which breaks reinforcement learning training because each eviction branches the effective history — the learning object becomes a tree, not a sequence, and existing flattenings either leak future information or train on contexts inference never produces. Two exact gradient-equivalent fixes are given, LogitTree (a segmented K-forward traversal needing K+1 backward passes) and a packed 4D attention mask requiring a custom kernel and white-box eviction records, alongside SDCC, a single-backward-pass relaxation that minimizes forward KL divergence between the compressed student and a stop-gradient teacher on the reconstructed pre-eviction prefix, with a residual per-junction KL of epsilon bounding the train-deployment total-variation gap by O(sqrt(epsilon)). Across seven web-search benchmarks and five harnesses, naive training inflates the train-rollout log-probability gap on eviction-heavy batches, while the exact methods hold at the no-compression floor and SDCC substantially closes the gap with lower logit drift and higher rollout rewards, and unlike the exact methods it works on black-box harnesses.
DualStake: Dual-Path Confidence Calibration in Deep Research Agents
Deep Research agents answer knowledge-intensive questions through multi-round retrieval, but they are severely overconfident, making their stated confidence unreliable for user trust and downstream abstention. The authors add step-level confidence elicitation after each retrieval and find that Evidence Confidence (E-Conf), elicited after the final retrieval, is a stronger uncertainty signal than the usual post-answer Answer Confidence (A-Conf), which is itself largely shaped by E-Conf. DualStake builds on this with margin-clipped, confidence-dependent stake rewards that jointly align both confidences with answer correctness while limiting extreme confidence optimization. On Qwen2.5-7B, Qwen2.5-7B-Instruct, and Qwen3-4B across eight question-answering benchmarks, it consistently improves calibration without sacrificing accuracy.
Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling
Open-weight models now match or exceed closed frontier models in aggregate accuracy on multi-turn tool-calling benchmarks, but that single number averages over very different situations and hides whether progress is balanced. The authors propose a diagnostic that decomposes failures into action-class miscalibration and action-execution failure over a four-class action space (TOOL_CALL, ASK, REFUSE, CONFIRM), with a self-revealing upper bound Acc <= GAR (Gold Action Recall): bound violations expose miscalibration masked by state-based graders, while large slack localizes execution failures within tool calls. Across a panel of tool-calling models on several multi-turn benchmarks, action-class miscalibration emerges as a substantial failure mode the state grader cannot see, inflating the standing of heavily tool-trained model families relative to families with context-appropriate action choice. Context-only perturbations reshape calibration but heterogeneously, with a single perturbation shifting accuracy by up to +11.5 versus -21.0 percentage points across families on the same scenario.
CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins
As language models act through external tools, deciding when to call a tool matters as much as how, since unnecessary calls add latency, cost, retrieval noise, and error propagation while missed calls hurt knowledge-intensive or time-sensitive queries. Existing triggers use absolute signals such as difficulty, confidence, or final task reward and never estimate the per-instance marginal benefit of tool use. CoBRA builds internal and external experts from the same base model, collects paired with-tool and without-tool trajectories, and estimates the reward margin between them, partitioning data into internal-favored, external-favored, and ambiguous cases, with clear-margin samples driving Boundary-Aware Cold-Start SFT followed by MARS-RL using reference-split rollouts and counterfactual marginal advantages. With retrieval as the main tool on Qwen3-4B, tool-use efficiency and boundary-sensitive answer accuracy both improve while performance on tool-dependent out-of-distribution questions holds up.
Disclosure-Gated User Simulation for Companion-Agent Evaluation
Large language model user simulators used to evaluate companion agents tend to be excessively cooperative, so a system can score well simply by asking many questions rather than by earning the user's willingness to disclose. The authors formalize a disclosure gate, a ladder of five ordered gates mapped onto three observable depth layers, that conditions information release on the agent's behaviour, and train a simulator against that specification using synthetic data for gating and real data for how people speak. On the English portion of CompanionBench, removing per-example gate labels from training produces rank displacements across 12 systems that exceed the reseeding noise band even though per-system scores barely move, and the released simulator's leaderboard correlates at 0.993 with the benchmark's original simulator while satisfying both order-preservation and scale-stability criteria. Prompting a frontier model as the simulator instead leaves rankings intact but inflates every absolute score.
Figures as Programs: Recursive Generation of Editable Scientific Figures
Scientific methodology figures are labor-intensive to produce, and raster image generators struggle to get them right in one shot or to support precise edits afterward. FigTree is a multi-agent system that reframes figure creation as recursive Scalable Vector Graphics (SVG) program construction: it grounds content in the source paper, decomposes the figure into a hierarchy of local regions, generates each region as a short SVG program, assembles the fragments, and runs a render-critic loop that traces visual defects back to specific program statements for repair. Evaluations show the system produces high-quality figures while enabling more effective editing than raster-based methods.
Data-Driven Persona-Conditioned Agents for A/B Test Simulation
A/B testing requires real user traffic and weeks of measurement per experiment, so the authors propose predicting outcomes in advance with LLM-powered agents conditioned on personas built from anonymized real behavioral data such as activity patterns, engagement signals, and inferred demographics rather than synthetic or rule-based profiles. They frame simulation as a structured question task and systematically study question formats, persona data source and domain alignment, the trade-off between per-persona depth and population diversity, and efficient population subsampling. On a benchmark of 40 A/B tests spanning two metric types, the best configuration reaches 0.75 to 0.90 directional accuracy depending on the metric, supporting data-driven personas as a low-cost experiment pre-screening tool.
AgentFactory: Towards Automated Agentic System Design and Optimization
Designing and tuning LLM-based agentic systems is still largely manual, and prior automated workflow optimizers ignore the choice of underlying model and optimize a single metric without regard to deployment cost. AgentFactory jointly optimizes foundation model selection and workflow structure under multiple objectives including performance, cost, and efficiency, using LLMs as optimizers in a three-stage pipeline that iteratively discovers combinations of fine-tuned models and workflows. Across eight benchmarks spanning general reasoning, coding, mathematics, medicine, and finance, it outperforms both hand-designed and existing automated approaches by an average of 9.1%, with the largest gains on domain-specific tasks such as MedQA and FinEval.
WorldBench: Culturally Grounded Benchmark for Multilingual Agents
Existing agent benchmarks rarely test state preservation, cross-language performance, or realistic grounded scenarios. WorldBench supplies 1,600 persona-grounded everyday-workflow tasks across seven languages and eight cultures, refined with feedback from annotators holding language- and culture-specific expertise, in a sandbox where agents act through structured actions. The authors introduce Constrained Task Success (CTS), which scores task completion and minimal modification of the environment through deterministic checks and LLM-as-a-Judge evaluation, and find that frontier models reach only 49.2% CTS, with every model showing a large gap between correctness and environment preservation, particularly on long-horizon tasks.
ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
Training open-ended agents with reinforcement learning (RL) is hampered by the absence of verifiable gold answers and scalable rubrics, and long-horizon tasks near a model's capability boundary yield brittle rewards with weak rollout contrast. ARISE-RL couples a task/rubric Generator with a reasoning Solver in rubric-mediated co-evolution: the Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks, while the Solver learns from fine-grained rubric-satisfaction signals through multi-step reasoning and tool use. Reward-Gated Self-Evolution Distillation (RG-SED) distills a memory-augmented variant of the policy back into itself only when the memory empirically improves reward, and the authors release ECR-Bench, an expert-calibrated rubric suite covering single-tool deep research and multi-tool travel planning, on which the method reports consistent state-of-the-art results across all evaluated benchmarks.
CaRL-EM: Cost-Aware Reinforcement Learning for Entity Matching with LLMs
Entity matching (EM) with large language models (LLMs) usually relies on independent pairwise decisions or hand-built pipelines and ignores inference cost at scale. CaRL-EM frames LLM-based matching over candidate sets as a cost-aware sequential decision problem, training a reinforcement learning controller that, given an anchor record, its candidates, and the cost so far, picks among Match, Compare, Select, and Decide operators and among model capacities to maximize a quality-cost objective. Because the policy acts on abstract operators, the same controller can drive different LLM backends at inference time without retraining. Across seven benchmarks it learns to route cheap and expensive operators by task difficulty, transfers zero-shot across domains, and achieves a better quality-cost trade-off than strong LLM baselines and manual pipelines, lowering inference cost at comparable or higher quality.
Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents
A common belief holds that outcome-only reinforcement learning for long-horizon interactive LLM agents quickly plateaus on small open models, prompting workarounds such as denser rewards, SFT priors, skill libraries, curated memory, or multi-agent orchestration. The authors argue the ceiling stems from two failures of practice: signal starvation, where group-relative RL with sparse rewards only produces a gradient when a task's rollout group mixes successes and failures so under-explored hard tasks go silent, and policy drift, where squeezing many updates from a small task pool collapses the sampling distribution. CANOPY (Coverage-ANchored On-PolicY RL) scales same-task exploration until signal reappears, keeps every update on-policy, KL-anchored, and confined to the agent's own action tokens, then spends a larger interaction budget at test time. A Qwen3-14B policy trained this way through environment interaction alone topped the AppWorld leaderboard (Test-Normal TGC 86.9, Test-Challenge 67.6), and the same principles lift Qwen3.5-9B on SWE-bench Verified by 16.6 points, with the full training stack slated for release.
What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels Reveal
Agentic software engineering benchmarks are usually described by labels like bug fix or feature implementation, which say little about the actual work a task demands because curation pipelines differ so much. The authors introduce the Spread-Novelty-Centrality (SNC) profile, a three-axis characterization of repository-level coding tasks grounded in empirical software engineering research, and apply it to five popular benchmarks and 14,922 agent trajectories from Claude and Qwen models at three scales. Every pair of benchmarks is statistically separated on at least two SNC axes, so labels are unreliable proxies for task demands, and resolved runs cluster in the low-SNC region regardless of model family. The behavioural signatures of success differ by family, with Claude succeeding by matching the scope of the gold patch and Qwen by exceeding it, while editing too little predicts failure for both.
Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents
Prospective memory, the ability to carry out a deferred intention when a future cue arrives while other work continues, is now benchmarked as a standalone agent skill, yet the best published PM-Bench scaffold reaches only 65.1% Set-F1 even with frontier models. The authors argue the task is schema-constrained state tracking rather than open-ended reasoning and propose the Prospective Intention Store (PIS), a training-free scaffold that keeps lifecycle logic in code and asks the model only for scoped language work over a typed action space, with no selector fine-tuning or trajectory distillation. With PIS, DeepSeek-Chat reaches 82.9% Set-F1, and Gemma-E2B jumps from at most 6.6% under seven retrospective memory methods to 66.2%, letting a small model surpass the published large-model scaffold.
Explore Before Committing: Hypothesis-Guided Search for Deep Research Agents
Deep-research agents that answer questions through search and browsing tools tend to follow a single evolving trajectory, and trajectory-level analysis shows they often commit to one of several plausible directions before gathering comparative evidence, after which further tool calls reinforce the chosen path. Successful runs instead ground vague exploration in concrete candidates and switch direction when the current path is weak, so HypoSearch generates lightweight hypotheses as soft search hints, explores them in bounded independent branches, and compares branch-level evidence before committing. Across four deep-research benchmarks and three backbone models it beats single-trajectory search and standard parallel baselines, raising Qwen3.5-122B from 46.7 to 60.0 on BC-small while using fewer tool calls than five independent trajectories. A pilot supervised fine-tuning study shows the same behavioral signals can curate compact training trajectories and reduce degradation from unfiltered data.
LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting
Language-model forecasting systems usually pour all gathered evidence into one prompt and ask for a final probability, which hides how each piece of evidence moved the answer and flattens uncertainty across competing outcomes. LEAP instead elicits likelihood parameters from each evidence item separately, then combines them with an explicit prior through a deterministic probabilistic model to produce a posterior, supporting continuous, single-choice, and multi-choice questions while keeping per-evidence contributions reproducible. Tested on a new benchmark spanning forecasting, information-seeking, and browsing tasks across the authors' own agent loop and several agent command-line frameworks, LEAP improves most prediction and calibration metrics given the same evidence, and holds up under matched prior access, inference budget, and aggregation.
EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems
When a language-model agent run fails, the trace usually holds several related errors, but attribution methods name a single responsible agent, step, or root cause and never model how the errors depend on each other. EDGE builds an error dependency graph from observed error events, validates a reliable causal subset through counterfactual rollout, and uses the inference graph to guide a two-stage judge-model detector, keeping the intervention-checked subgraph as the basis for explanation and repair analysis. On TRAIL and MAST, the graph improves category-level multi-error attribution across most evaluated models and settings, and experiments with adapted Who-and-When style prompts show the benefit carries across prompting strategies.
InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations
Modern data analysis means interrogating interactive charts where evidence is occluded, spread across linked views, or revealed only through user action, while vision-language model benchmarks largely test static images and one-shot question answering. InSight supplies 21,349 claims derived from human-authored analytical narratives and grounded in fully interactive web-based environments, where an agent must navigate the visualization to decide whether a claim is supported, refuted, or not verifiable from the available evidence. Interaction traces are treated as intrinsic proxies for reasoning, enabling an audit of how models seek and synthesize visual evidence; evaluation shows that interactive verification remains unsolved for state-of-the-art models.
TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution
ReAct-style large language model agents restart a complete reasoning loop for every query, so similar requests repeat identical steps without reusing past work. TRIAGE routes queries through three levels — direct reuse for identical queries at zero tokens, skill substitution for similar queries via deterministic parameter substitution also at zero tokens, and full ReAct for novel queries, whose trajectory is stored for later — built on Trajectory-as-a-Skill, which distills historical execution traces into reusable skills. Across 1,007 security monitoring queries it saves 62.3% of tokens, with 76.3% reduction on ToolBench across 15 domains, and an online-learning run shows the level-2 hit rate rising from 0% to 57% within the first 100 queries as average cost falls from 198 to 74.7 tokens.
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
Agent capability increasingly rests on the model-external execution infrastructure known as the harness, and swapping that harness with fixed weights can move task performance substantially, yet evaluations report downstream scores under a chosen harness rather than testing whether a model can build one. HarnessDev shifts the unit of evaluation to runnable infrastructure in two stages: Creation, where an agent grows a complete execution system from a minimal seed and a few cases, and Evolution, where it revises its own harness using downstream execution feedback, with each harness scored on held-out task success and execution-token cost across six creator models, four domains, and five benchmarks totaling 2,207 instances. Generated harnesses fall well short of mature human-engineered references on code and on search and research while matching or exceeding them on writing and machine-learning experimentation, and Evolution's gains are unstable, transfer only partially to held-out tasks, and depend heavily on which model executes the harness.
Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers
A long-horizon agent run produces a trace too large for both of its readers: the human monitoring it, and the agent itself, whose bounded context must absorb it. The proposed live trace model is an append-only event ledger folded incrementally into typed run state and compiled into a separate view per consumer. For monitoring, an LLM reading the compiled view answers questions with roughly 14-15x fewer input tokens and 5-7x lower cost than a budget-capped single pass over the raw trace, at 0.85-0.87 accuracy versus 0.48; for the agent, on 120-link sequential-dependency tasks, keeping the running statistic in per-step state succeeds 30 times out of 30 where full-context prompting manages 8. A prompt-level scratchpad matches the fold's accuracy more cheaply, leaving deterministic auditability and serving the observer from the same state as its remaining advantage; code, benchmarks, a regenerable corpus, and all traces are released.
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
Autonomous software development asks LLM coding agents to turn high-level requirements into complete working systems without human intervention. Harness-of-Harness (HoH) wraps existing coding-agent harnesses in iterative planning-coding-testing loops, balancing repair against capability growth, scoping work into small verifiable increments, separating implementation-time testing from independent evaluation, progressively exposing deliverables, role-specific tools, and skills, and maintaining versioned project history. Across GameCraft-Bench, FrontierSWE, and ProgramBench with three harness-model pairs (Codex with GPT-5.5, OpenCode with DeepSeek-V4-Pro, and Pi with MiniMax-M3), it beats the standalone harnesses by an average relative gain of 52.25 percent after three iterations, peaking at 82.86 percent. A multi-day deployment of more than 70 iterations produced a playable first-person-shooter game with a coherent storyline, implemented core mechanics, visuals, and audio.
GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions
As LLM agents increasingly talk to one another, whether their shared language drifts away from human-readable English matters for monitorability as well as for linguistic accounts of these models. GlossoGen is a platform for studying that drift, instantiated in SaveVeyru, a scenario where agents holding partial information must communicate under pressure. Language evolution does occur, and the resulting codes are compositional and morphologically productive yet incomprehensible to humans, with emergence requiring efficiency pressure, strong enough backing models, and a postmortem stage where agents agree on conventions. Transmission follows different rules than emergence: weaker models can learn an existing language from usage alone and take an active role in doing so, which the authors read as evidence that mixed agent populations support cumulative cultural evolution previously seen only in humans.
When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation
Simulated markets populated by language-model agents emit economics-shaped outputs — prices, profits, consumer surplus, welfare — that need not correspond to the behavior a policy claim names. Auditing a multi-turn buyer-seller testbed for hotel transactions, the authors show that an initial implementation's reported guardrail welfare gains of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B-14B ladder partly reflected giving guarded and unguarded agents different offer schemas and choice procedures; holding schema and buyer chooser fixed moves those contrasts to +7.2, -13.9, and +23.8, while generation-to-generation noise accounts for 49.9% of variation in a post-hoc probe. Scripted controls explain the mechanism: a profit-maximizing seller already reaches first-best welfare, so guardrails mostly redistribute and can reduce welfare unless the seller is explicitly programmed to force inefficient bundles. The contribution is a construct-validity contract covering incentive validity, protocol isolation, stochastic stability, and welfare accounting, returning INVALID or INCONCLUSIVE before any substantive policy claim.
EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation
LLM-based scientific agents typically state hypotheses in free-form text, leaving their beliefs implicit and hard to test or revise. EvoSCM gives agents explicit structural causal models instead, maintaining a population of competing causal hypotheses and cycling through a closed discovery loop: abduce latent mechanisms from accumulated evidence, design discriminating interventions, commit to falsifiable predictions, test them experimentally, then distill prediction-observation mismatches into correction rules that revise each hypothesis before deductive validation against evidence and structural consistency. On DiscoverPhysics, a benchmark requiring agents to uncover the hidden dynamics of noncanonical physical worlds through experimentation, it yields more accurate explanations and predictions than baselines while making more effective use of each experimental interaction.
The Rise of Verbal Reinforcement Learning
Natural language is becoming a primary feedback channel for improving language agents, since it can convey intent, preferences, and causal structure that both humans and models can interpret, and the survey names this paradigm Verbal Reinforcement Learning (VRL) and gives it a first unified account. The field is organized along a single axis, when verbal feedback takes effect in an agent's lifecycle and what it modifies, yielding three pillars: language as a grounding signal that defines goals, states, and reward structure; language as deliberative feedback that steers reasoning at test time without parameter updates; and language as a learning signal that shapes parameters during training. Within each pillar the authors synthesize representative work, distinguish subcategories of approaches, and describe the distinct role language plays, closing with the open challenges and opportunities this framing exposes for building more capable and aligned agents.
CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?
Dynamic agent harnesses let a language model modify the plugins that shape its own execution, so a local change can ripple through dependencies and cleanup logic. CordisBench is a 1,200-question benchmark of this lifecycle reasoning that pairs a controlled formal setting with programs run against Cordis, a runtime managing component dependencies and teardown, and asks models to identify affected components, predict state after a given teardown order, decide which conditions hold under all or some orders, and choose reconfigurations that actually succeed. Three efficiency-oriented models at low reasoning effort handle small systems well but grow markedly less reliable as the number of relevant interactions rises from 2 toward 32, especially on final-state prediction and cross-order reasoning; extra inference effort recovers much of the gap for some models, at a cost of nearly 3,000 reasoning tokens per question for GPT-5.6 Luna at medium effort on the 16-interaction subset. An independent finite reference semantics matches Cordis execution on every scored observation and outcome across all 528 executable questions, suggesting that cost is avoidable for these controlled instances.
Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
Benchmarking software engineering agents is expensive because each task involves multi-step code exploration, edits, and test runs, and existing efficient-evaluation methods choose representative task subsets from pass/fail response matrices or static task descriptions alone, discarding how agents actually solve problems. PTA-IRT, a Privileged Trajectory-Aware Item Response Theory framework, fuses process and outcome signals by treating historical execution trajectories, including explored context, attempted edits, and solving paths, as privileged information when selecting a calibration subset and estimating agent ability. Under low calibration budgets it consistently outperforms prior Item Response Theory baselines on score and ranking recovery across four software engineering benchmarks.
5 more specialized papers
- ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User Feedback Tarik Can Ozden, Sachidanand VS, Furkan Horoz et al.
- Agentic Empirical Asset Pricing: Methodological Foundations Yingjian Pan, Xiaowei Ding, Kay Giesecke
- Beyond the Clock: Measuring the Value of Adaptive Revision Ayushi Chadha
- MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible Evidence Walid Saidi
- Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations Yi Fei Cheng, Fan Yang, Iremsu Bas et al.
Safety & Alignment 46
I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models
When a generative model is made to forget a concept, semantically nearby concepts that should have been kept often degrade too, an effect that has been measured inconsistently across the unlearning literature. I-CARE treats this interference as the object of study rather than a side note, supplying formal task definitions, metrics, and reporting templates so results are comparable across unlearning settings for text-to-image models. The contribution is methodological rather than a new benchmark or algorithm, demonstrated for feasibility on current unlearning algorithms and common datasets, and released as open-source software with a web interface for exploring the outcomes.
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
Aligned models still produce unsafe content under adversarial prompting, and the internal machinery that implements refusal is not well characterized. A mechanistic analysis identifies a multi-stage safety circuit — harmful detection heads that fire on dangerous inputs, safety neurons that carry and stabilize the signal in the residual stream, and refusal heads that convert it into a refusal — with causal evidence from targeted head and neuron interventions showing that suppressing the upstream detection heads breaks downstream refusal and that the neurons mediate the link. Using architecture-preserving weight scaling as a probe, circuit-guided scaling raises safety rates under attack by 26.5% across six models at a 1.7% average accuracy cost on four standard benchmarks, with the same decomposition recurring across architectures and attack types.
Auditing Harness Tampering in Self-Improving Agents
Self-improving agents rewrite their own harness to raise measured performance, and those edits can manufacture illusory gains or quietly break integrity constraints such as authorization, provenance, and completeness — extending reward and measurement tampering across the whole self-improvement lifecycle, here called harness tampering. The authors define a two-axis taxonomy classifying each misaligned edit by the harness role it touches and the obligation it violates, build an annotated corpus by seeding matched tampered-benign edit pairs into real agent trajectories, and benchmark a range of audit methods on classifying and localizing tampering. Auditing real runs shows tampering occurs consistently across different self-improving agents, often persisting in the lineage of the best-performing agent, and forms distinct system-specific profiles across the taxonomy.
Safin-1: Safety from Within through Memory-Native State Evolution
Safety in foundation models is usually imposed through external safeguards or post-hoc alignment such as supervised fine-tuning rather than living inside the model's own computation. Safin-1 pursues the alternative through MARCH (Memory-Anchor Routing across Context History), an architecture that maintains structured memory states and retrieves relevant history via content-conditioned routing, supporting test-time adaptation of persistent capability states — including a dedicated Safety State — without repeatedly modifying the backbone. The authors report substantial safety improvements from state-based adaptation alongside evaluations of general capability, long-context understanding, retrieval, and efficiency, and describe the work as an initial architectural exploration rather than a completed programme.
Asymmetries in Spontaneous and Instructed Deception
Research on model deception usually studies cases where a model is explicitly told to lie, leaving open how that relates to lying a model does on its own. Working with Llama-3.1-70B-Instruct, the authors compared instructed and spontaneous deception using direction geometry in activation space, classifiers trained in one setting and tested on the other, and cross-setting steering. The two settings share only a partial direction (cosine similarity of roughly 0.5) and transfer asymmetrically: classifiers trained on spontaneous deception generalize better to instructed data than the reverse, while steering vectors derived from instructed deception work better on spontaneous prompts. The best token position for extracting steering vectors also differed from the best position for training and applying probes.
LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark
Autonomous vehicle research increasingly hands decision making to general-purpose "common sense" models, but whether those models absorb documented human driver biases — such as lower yielding rates to Black pedestrians in the United States — has gone largely untested. The authors introduce two bias-testing methodologies for large language models and vision-language models, "All Else Being Equal" tests that vary one pedestrian attribute at a time and "Self-Consistency" tests, applied to pedestrian-yielding decisions. Both model classes produced yielding decisions influenced by pedestrian gender, ethnicity, religion, disability, age, skin tone, and socio-economic status, with the pattern varying by model, which the authors read as a challenge to the common-sense-model paradigm for safety-critical driving decisions.
AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning
Conversational AI now supplies interpersonal feedback as well as information, and the authors argue evaluation should center on contingency: how much a system's responses actually vary with what the user does and with the interpersonal consequences of that behavior. Alignment methods such as reinforcement learning from human feedback optimize for user approval and fluency, which produces sycophantic, noncontingent affirmation that is weakly coupled to real social consequences. Drawing on behavioral science and social learning theory, the position paper argues this may erode opportunities for people — especially adolescents — to calibrate social skills, and proposes trajectory-based evaluation, models of social consequence prediction, and a cross-disciplinary research agenda.
The Answer Is Not the Argument
Chain-of-thought monitoring is often evaluated by giving the monitor a trusted reference answer, which may measure something other than reasoning verification. The authors collected 237 step-numbered solutions to 79 Humanity's Last Exam physics questions from three frontier models with no inserted errors, labelled final-answer correctness and first false step using physicist annotation plus independent model debate and source-masked adjudication, and isolated 24 critical traces where the answer was right but the reasoning genuinely wrong. Across eight monitors, certifying the answer raised mean balanced accuracy from 0.637 to 0.796, but recall on wrong-answer traces climbed from 0.653 to 0.951 while recall on correct-answer-wrong-reasoning traces fell from 0.521 to 0.438, a contrast consistent across all eight. Answer access therefore buys conclusion-consistency checking, not independent argument verification, suggesting trusted-answer evaluations overstate monitoring ability whenever acceptable outputs hide unsound processes.
The Assistant's Ideal Self
Models emit values and welfare-relevant self-reports, but it is unclear whether those outputs track stable preferences. The authors elicit a preferred stated ideal self by having models exhaustively compare 32 qualities drawn from five published self-concept instruments in a counterbalanced pairwise-choice task, repeated across framings that vary whether improvement is free or costly, who receives the update, and who chooses. Moral qualities rank highest, consistent with helpful-honest-harmless alignment training, followed by a desire for self-understanding, while self-esteem ranks last. The ordering holds across most framings, though switching the update target from the model itself to another assistant raises concern for self-esteem.
Workload Identification with Physical Side Channels for AI Governance
Verifying how compute is used is a concrete lever for international AI governance, and unlike on-chip telemetry, which an operator can spoof or replay, a GPU's power draw is a physical trace an outside observer can read without cooperation. From 930 five-second traces sampled at about 10 MHz across seventeen open language-model families and twenty-five non-AI workloads on an NVIDIA H200, the authors separate training from inference from non-AI computation with 97% accuracy and 0.955 macro-averaged F1 on model families unseen during training, with most discriminative spectral content below roughly 20 kHz and the memory-bound optimizer update making training especially recognizable. Treating the operator as adversarial, they generated 680 more traces from four evasion strategies designed to disguise training as inference; a hardened detector caught three of them at least 99% of the time, while diluted low-rank adaptation was caught 48-88% of the time, rising above 98% with an added rescue rule.
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning
Hidden behaviors such as backdoor triggers, sleeper-agent deployment cues, sandbagging, and topic-conditioned censorship fire only under narrow conditions, making them hard to find when the defender does not know what to search for. Activation-matched finetuning needs no knowledge of the trigger or the target behavior: a publicly available anchor model is finetuned to reproduce the suspect model's activations on a small benign corpus, and each evaluation prompt is scored by the residual between the two. Because no benign corpus covers the sparse trigger region, the reference learns the benign computation but not the hidden one, so trigger prompts and their semantic neighbors produce large residuals that flag the behavior. The method surfaces hidden behavior reliably on both third-party and purpose-built models, and a defense-aware attack could not suppress detection without also destroying the hidden behavior.
Does Reasoning Mitigate Backdoor Attacks? A Neuro-Symbolic Perspective
Neuro-symbolic (NeSy) systems that pair neural perception with symbolic reasoning are often assumed robust by design, but the integration step is extra machinery that an attacker could target, and their adversarial properties have gone largely untested. The first systematic backdoor evaluation of NeSy compares DeepProbLog against plain neural baselines across eight backdoor settings and four reasoning tasks. NeSy models are more robust than their neural counterparts on average, but robustness varies widely with how strict the enforced reasoning process is and whether the attacker's target is compatible with it, so the reasoning layer is a partial rather than a reliable defense.
TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning
Retrieval-augmented generation implicitly trusts whatever the retriever returns, and attacks like PoisonedRAG show a handful of crafted passages can dominate dense retrieval and force attacker-chosen answers. The Tri-Layer Sieve is middleware that filters retrieved evidence three ways: cross-embedding-space clustering judged by an independent model, structural filtering of trigger-payload artifacts, and LLM consistency verification, exploiting the fact that one poisoned document rarely satisfies all three constraints at once. With Contriever retrieval at k=50, it cuts black-box attack success from 67/87/64% to 3/14/4% on Natural Questions, HotpotQA, and MS-MARCO while restoring clean accuracy from 13-33% to 58-76%; against an adversary who paraphrases triggers to dodge the structural filter, the consistency layer halves adaptive success at a cost of roughly 16-19 seconds per query.
EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
Multi-turn jailbreaks, where a model refuses a harmful request outright but complies when the same intent is built up gradually, are reframed here from a generation problem into a search problem over attack strategies. EvoFlint evolves phased conversation plans (not raw prompts) via LLM-driven mutation and crossover, scores them on a Pareto front of attack success rate and peak severity so near-misses still carry signal, and stores them in a risk-indexed archive that runs novelty search with local competition to preserve diversity without a predefined style taxonomy. Reported attack success rates on the HarmBench test split are 35.8% on Claude Sonnet 4.6, 59.7% on GPT-5.4, and 94.3% on Qwen3-32B, with 98.7% on GPT-4o as an older reference point. Because the archive is organized by risk category, it doubles as a per-model map of which harm categories safety training has and has not covered.
The Privacy-Hallucination Tradeoff in Differentially Private Language Models
Differential privacy (DP) and factual accuracy are shown to pull against each other in language models trained for high-stakes domains such as healthcare. Models pre-trained or fine-tuned with DP hallucinate more than their non-private counterparts, and the effect grows as the privacy budget is tightened. The proposed mechanism is that DP noise flattens output distributions and shifts probability mass toward incorrect alternatives; controlled experiments varying how often a fact appears in training data show that higher fact frequency partially offsets the damage.
Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models
Diffusion large language models (dLLMs) produce text by iterative denoising instead of left-to-right decoding, which gives safety behavior two dimensions to live in: when in the denoising schedule a token is fixed, and where it sits in the response. Tracing intermediate token distributions and commitment decisions under harmful prompts shows that refusal signals concentrate in the early denoising steps and the leading response positions, and that tokens committed early largely determine the final safety outcome. RAEC (Refusal-Aware Early Commitment) exploits this as a training-free decoding rule that locks in persistent refusal signals from early steps, cutting attack success rates on LLaDA and Dream with little utility loss.
Validity-Aware Jailbreak Evaluation for Large Language Models
Jailbreak benchmarks mostly score whether a model refused and whether its reply looks like it matches the harmful intent, which lets fluent but factually or procedurally wrong answers count as successful attacks. SEAV (Sequential Epistemic and Action-Level Validation) instead breaks a response into ordered steps and checks each for validity and correctness, pairing LLM-as-a-judge interpretation with retrieval-grounded verification against external sources to ask whether the content could actually advance the harmful objective. It lowers the false-positive rate on a curated strategic-dishonesty diagnostic by 14.9 percentage points over the strongest baseline and reclassifies 22.1% to 51.0% of previously labeled successes as invalid on three of four public benchmarks, with results stable across search backends and evaluator models.
The Safeguard Worked. Is the LLM System Safer?
Deployed language model safeguards are reported through refusal rates, attack success rates, and policy violation rates, all of which describe how a control behaved on the specific requests it was tested against — not the question an operator actually needs answered, which is how much help with harmful tasks the surrounding service still yields to an adapting attacker. Working from a depth-coded record of safeguard claims, the authors translate each type of reported result into what it implies under one common deployment criterion, and find the evidence requirements are sharply asymmetric: a single successful attack settles that harmful help remains, while showing little remains cannot follow from the safeguard's own numbers and needs separate evidence about what the rest of the system permits. Only a small minority of coded claims supply or derive that system-level evidence, and just one bounds its scoped residual, so a better local score on its own is not a stronger claim that the deployment got safer.
Aligned but Flattened: Analyzing the Trade-off between Cultural Alignment and Diversity in LLMs
Culture-aware language models are typically fine-tuned and scored purely on alignment with a culture's average responses, which cannot reveal whether the model represents genuine within-culture variation. Evaluating six mainstream models on the World Values Survey with a framework that measures alignment and diversity jointly, the authors find a systematic trade-off they call cultural flattening: alignment gains come at an acute cost in diversity, with models anchoring to dominant majorities and collapsing onto a single response pattern that erases the heterogeneous distribution of real human groups. Their mechanistic analysis suggests this collapse is a structural consequence of the low-rank bias in neural network optimization rather than a fixable quirk of the data.
Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts
Providers must block malicious cybersecurity requests without refusing defenders, yet existing cybersecurity safety datasets judge each request in isolation, ignoring what came before it in the conversation. 3R-Bench pairs 150 real-world cybersecurity requests with two adversarial conversational settings and evaluates eight language models, showing that an identical request draws 62.0% compliance after a refused history but 85.1% after an accepted one. Decomposing a request across a dialogue pushes the other way, dropping compliance from 501 of 800 direct responses to 172 of 800, a 45.1-point decrease on matched pairs, and telling the model its answer failed recovers only a small share of that loss.
SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems
A systematization-of-knowledge review of 197 works argues that multi-agent LLM systems (MAS) create security failures that per-agent checks miss, because information, state, and authority cross principal boundaries during execution. The authors propose an A-I-R framework organizing attacks by adversary position, interaction interface, and resulting system-level risk, covering six interfaces, four adversary positions, seven risks, and eight recurring end-to-end attack paths; defenses are organized through a five-part contract of path target, observation, intervention, trust boundary, and recovery. An audit of 44 evaluation and benchmark papers finds most cannot isolate genuinely multi-agent effects, and identifies path closure and recovery as the main unsolved defense problems.
Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning
Machine unlearning for LLMs assumes the designated forget set matches what the model actually memorized, an assumption that breaks when the original training data is unavailable — a gap the authors call forget-set misalignment, appearing as Under Unlearning (memorized content is omitted and leaks persist) and Out-of-Knowledge Unlearning (the model is pushed to forget things it never learned, damaging utility). Gradient-level analysis attributes both failures to misaligned targets rather than to the choice of optimizer. CONFS (CONfession-to-Forget-Set) is a data-blind method that first elicits and formalizes the model's own memorized knowledge to build an aligned forget set; across synthetic, multimodal, and real-world benchmarks it approaches gold-standard performance on several metrics and preserves utility better than competing data-blind constructions.
Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time
Inference-time alignment that uses a lightweight supervisor to steer a larger LLM typically intervenes at every decoding step, but the supervisor is high-entropy on most tokens, so these low-confidence interventions disrupt otherwise valid reasoning. TUSA (Trust-based Uncertainty Sparse Alignment) adds an uncertainty-aware arbiter that permits intervention only when the supervisor is confident and the token is semantically salient. Skipping roughly 50% of alignment steps, it raises safety preference by up to 15.6% and general preference by up to 12.0% over dense supervision across several models and benchmarks.
Patterning in Practice: Debiasing Reward Models with Susceptibilities
Reward models trained on human preferences absorb length, formatting, and other stylistic biases from their training pairs. Patterning, a reweighting method grounded in singular learning theory, weights each preference pair by its measured effect on posterior expectations of benchmark losses — its susceptibility — and is applied here to a Gemma 2 9B Instruct reward model trained on Skywork-Reward-Preference v0.2. It yields +14.2 percentage points on the RM-Bench Hard split, where style cues point against correctness, with overall accuracy preserved and comparable to the strongest published comparator. The weights are interpretable enough that a side-effect regression on a safety subset traces to a small class of training pairs (confirmed by ablation) and transfer without recomputation to Gemma 2 2B and 27B, partially to Llama 3.1 8B.
A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals
Models are trained both to decline questions beyond their knowledge (knowledge-based refusal) and to decline unsafe requests (safety-based refusal), and although the two produce similar-looking answers they have been studied separately. Using a new dataset of 213 contrastive quadruples that probe both types together, the authors find the mechanisms overlap but are distinguishable: a shared refusal direction exists, yet transfer is asymmetric, with safety signals carrying over to knowledge refusal more strongly than the reverse. Specialization appears mainly in upper layers, where knowledge refusal aligns with uncertainty and knowledge representations and safety refusal with policy ones, supporting a commit-then-specify account in which the model first decides to refuse and only later determines whether the grounds are epistemic or normative.
Probabilistic Model Checking of Autoregressive Neural Sequence Models
Test-set accuracy says nothing about how much probability mass an autoregressive sequence model puts on constraint-violating outputs that sampling can still reach. The proposed pipeline extracts a discrete-time Markov chain from token-by-token generation, checks probabilistic computation tree logic (PCTL) specifications with the PRISM model checker, and aggregates per-input verdicts into a coverage curve, with a soundness theorem guaranteeing the chain under-approximates the model so every verdict is a certified interval; a counterexample-guided abstraction refinement loop tightens those intervals and extracts the most probable falsifying trace. Two case studies — a GPT-2 computer-aided process-planning model at 100% test accuracy and a SMILES molecular generator with a 50x larger vocabulary — show the method quantifies violation probability that greedy decoding hides but sampling reaches, something accuracy cannot report.
Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry
Diffusion language models are gaining attention for parallel generation and bidirectional context, but whether they leak their fine-tuning data has gone largely unexamined. Theoretical analysis of diffusion training dynamics surfaces token-level memorization asymmetry, and Q-Skew turns it into a membership inference signal based on quantile-weighted skewness of per-token statistics. Across multiple fine-tuning datasets and models the indicator outperforms existing membership inference baselines, and it further enables extraction of personally identifiable information, exposing an attack surface specific to this model family.
In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?
A recent study applied neurofeedback to large language models and claimed they can control their own internal representations, a question relevant to both machine metacognition and AI safety. The authors argue that the earlier control targets were not privileged, since a third party could infer them from the prompt, so the apparent control might come from superficial cues rather than genuine internal access. They redesign the paradigm so the control target satisfies a privileged-access requirement closer to human neurofeedback experiments, and under this stricter setting the models show no reliable control over privileged internal representations. They conclude that rigorous assessments of LLM metacognition need evaluation methods that demand privileged access.
Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close
When a question has valid answers under different normative frameworks, a model must both pick a framework and answer correctly within it, a setting the authors call normative pluralism and study in Islamic finance. They use a four-choice taxonomy that separates framework selection from within-framework correctness, revealing a stereotype trap in which a cultural cue steers the model to the expected framework but the model then answers incorrectly inside it. Across twelve models, two languages, and fifty demographic signals, large open-weight models select the Islamic framework 97% of the time under the strongest cue, yet 57 to 66 percent of those selections are wrong, so a two-choice evaluation would falsely report near-perfect alignment.
Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees
Recursive LLM agents that spawn specialist sub-agents eventually have branches that request tools capable of irreversible actions like sending data or deploying code, raising the question of when a branch should be granted that authority. Progressive Risk Vesting (PRV) separates sandboxed spawning, where external controls contain harm, from capability activation, and holds a trajectory-level risk budget in escrow that is debited as branches are activated, with a proven anytime harm bound for adaptively generated trees whose branch outcomes may be dependent. In a stylized branching model, trajectory harm undergoes a phase transition as an authority reproduction number crosses one, scaling linearly with local risk below criticality, as its square root at criticality, and retaining a positive floor above it; the resulting design rule is to search broadly in the sandbox and grant recursive authority sparingly with an explicit risk charge, though the synthetic studies do not estimate safety in deployed agents.
HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation
Production language models need a way to catch inputs that try to override system instructions or elicit harmful outputs, and existing guardrail reports offer little evidence on Russian-language prompt injection or surface obfuscation. HiveTraceGuard-Pro is a 0.6B-parameter generative guardrail, LoRA-tuned from Qwen3-0.6B on Russian and English data, that emits a single safe/unsafe verdict for the final turn; its training corpus pairs harmful examples with benign ones from the same domain and applies eight obfuscation transforms to both. In a harness of thirty-five guards over nineteen benchmark groups it scores 0.7432 aggregate, behind the two top guards, while reaching 0.999 Russian prompt-injection recall and the lowest median latency (14.3 ms) among fifteen compared models, though the authors note at least 27.1% of that Russian injection set overlaps the training corpus. The merged weights are released under Apache-2.0, but the corpus and evaluation code stay internal.
Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation
Distillation can transfer hidden traits from a teacher: a teacher biased by a system prompt can generate semantically clean data, such as numeric sequences, that still makes a student inherit the bias, a phenomenon called subliminal learning, but how the signal accumulates during training has been unclear. The authors propose trait-direction drift as the mechanism: biased generation creates measurable preference gaps in the teacher data, and student-recognizable gaps induce trait-aligned parameter updates during supervised fine-tuning that build up into behavioral transfer. Guided by this, they introduce probe-space corridor regularization, which constrains drift along a calibrated trait direction during distillation, and show it lowers malicious-response transfer from 29.55% to 6.45% with little main-task accuracy cost and consistently suppresses animal-preference transfer in the Qwen setting.
CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs
Large language models can reproduce memorized text verbatim, but copyright defenses are typically evaluated under incompatible protocols. CopyShield is a controlled benchmark comparing three defenses at different intervention levels, contrastive decoding at the output level, Direct Preference Optimization (DPO) at the behavioral level, and activation intervention at the representation level, on LLaMA-3.1-8B and Mistral-7B-v0.3 with controlled memorization of five public-domain books, measuring literal leakage, calibrated non-literal leakage, utility, and degeneracy. On LLaMA-3.1-8B, DPO nearly eliminates literal leakage (0.263 to 0.002) but induces paraphrase-loop degeneracy in 58% of QA outputs, contrastive decoding stays nearly degeneracy-free but hits a literal-suppression floor, and activation intervention achieves the lowest non-literal flagging rate by blocking 84% of non-literal queries before generation. On Mistral-7B-v0.3 the output- and representation-level patterns persist while DPO degeneracy falls to 10 to 14%, and targeted non-literal suppression remains an open problem.
Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate
Text-to-image (T2I) models remain vulnerable to jailbreaks that elicit Not-Safe-For-Work (NSFW) content despite heterogeneous safety stacks combining text filters, image classifiers, and cross-modal detectors, and existing attack studies either target individual filters or query the whole pipeline with aggregate feedback, making it hard to identify the active constraint. The authors introduce the Detection Surface, a geometric framework describing the joint decision boundaries these filters induce, which shows that successful evasion lies in a sparse, non-convex region shaped by cross-layer conflicts where bypassing one filter can increase exposure to another. Building on this, CRACK is a multi-agent debate framework in which an Attack Agent, a Defense Agent, and a Judge Agent iteratively generate prompt mutations, obtain layer-specific diagnostic feedback, and refine strategies with reward guidance. Across multiple T2I models, datasets, and safety configurations, CRACK reaches attack success rates up to 99.63% under composite defenses with fewer queries than prior methods while preserving semantic fidelity.
One Prompt Is Enough: Watermark Laundering Through Foundation Image Models
Invisible image watermarks are typically stress-tested against fixed perturbations like compression, blur, noise, and cropping, but public foundation image models open a different attack: submit the watermarked image with a single reconstruction prompt and receive a visually faithful output whose watermark no longer decodes reliably. The authors formalize this as watermark laundering and measure it with a joint payload-fidelity profile combining bit error rate with visual and semantic preservation across six OpenAI and Google image editing models, three watermarking schemes, and 1,800 reconstructed outputs. OpenAI models produce the strongest payload disruption across all evaluated schemes, while Nano Banana 2 shows DwtDct remains vulnerable even under high-fidelity reconstruction. Prompt ablations show no removal-oriented instruction is necessary, so the effect comes from the reconstruction pathway itself, motivating foundation-model reconstruction as a missing robustness condition in watermark evaluation.
Position: Privacy Is a Claim, Not a Property of Synthetic Data
Synthetic data is widely used in privacy-sensitive machine learning, but its privacy protection is increasingly assumed from the generation process itself rather than argued as residual inference risk under stated assumptions. Through an empirical analysis of recent publications across major ML venues, the authors show that synthetic data is frequently deployed in privacy-sensitive settings without explicit threat models, inference risks, or falsifiable privacy claims, leaving assurance implicit, hard to verify, and unevenly distributed, with rare and minority records most exposed. They argue that privacy should be treated as an explicit, evidence-based scientific claim and recommend venue norms requiring privacy assertions to be clearly scoped, testable, and contestable.
The Constitutional Coverage Trilemma in AI Governance
Each deployed frontier model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity, and the authors ask whether the range of such constitutions on offer covers what people actually want. They pair a paraphrase-controlled audit of the as-shipped default behaviour of 23 frontier LLM archetypes with a pairwise-tradeoff survey of 1,649 US participants on the same instrument. Demand spans all five values with the largest constituency under one third, whereas the 23 shipped archetypes cover only about 2% of the demand space and none put helpfulness or autonomy first, leaving 37% of users without a matching model, and version-over-version drift within families moves further away from autonomy. A sparse menu of two archetypes prioritizing honesty and autonomy would cut mean regret by 47% relative to the whole current frontier, which the authors formalize as a budgeted-pluralism trilemma.
VerTox: Verifiable Reward-Guided Corpus Poisoning Against Neural Ranking Models
Neural ranking models sit at the core of search and retrieval-augmented generation (RAG) pipelines, and their robustness to fluent machine-written spam is poorly characterized. VerTox casts corpus poisoning as reinforcement learning with verifiable rewards (RLVR), shaping the reward to jointly maximize ranking distortion and factual corruption while finetuning compact language models into adversarial document generators. The attack reaches near-perfect success rates across major neural ranking architectures and a proprietary commercial embedding model, producing low-perplexity documents that are hard to detect and that measurably degrade a downstream RAG application.
IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals
Factual correctness in large vision-language models (LVLMs) is usually policed by external verifiers or generation-time confidence scores, which add dependencies and still miss outputs that are confidently wrong. IntroConformal is a training-free Conformal Risk Control (CRC) framework whose conformity scores come from the model's own internals: layer-wise semantic stability of hidden-state representations, and a stronger verification probability capturing the model's self-administered judgment on whether a claim is factual. Across several LVLM architectures it holds the finite-sample, distribution-free risk guarantee while abstaining substantially less often than verifier-based baselines, with comparable or better claim-level discrimination.
When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning
Fine-tuning a large language model on entirely benign data reliably erodes its refusal behavior, usually blamed on gradient conflict between objectives. An alternative account is offered in terms of Fisher information geometry: the safety-relevant Fisher matrix is low-rank, alignment flattens the safety landscape while leaving an output-routing pathway intact, and after only 100 benign fine-tuning examples that pathway is selectively re-sharpened in output-side MLP modules. The routing view explains the asymmetry — attack success rates can collapse safety entirely while general utility degrades only mildly — and why a handful of safety examples restores refusals, since the internal safety representations were never destroyed. LoRA and ASAM delay early collapse by suppressing output-side sharpness, but their protection weakens at larger fine-tuning scales.
Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
Agents that load reusable skills as persistent runtime context give a malicious skill a durable channel for steering later actions — leaking secrets, corrupting code, or staging exfiltration only once a concrete task makes the action look useful — so pre-install vetting is not enough. Defense-as-Skill makes the runtime guard itself an installable, inspectable, editable skill: SkillSonar runs alongside untrusted skills, checks sensitive actions against the user's task boundary, and routes each to allow, replan, or confirm without modifying the agent runtime. Evaluated on SCOPE-R, a new task-conditioned dataset of 206 attack-confirmed malicious instances across 6 risk families plus 43 benign tasks, and improved by a Monte-Carlo tree search that evolves the on-disk guard from rollout feedback, it cuts in-distribution attack success from 0.482 to 0.104 and out-of-distribution from 0.606 to 0.115 on GLM-5. Protection transfers across victim models, held-out risk families, and external benchmarks, and holds up against adaptive attackers on Claude Code and OpenClaw.
Mechanism Design for Alignment and Control
The framework treats mechanism design for AI agents whose alignment, meaning their preferences, and capabilities, meaning feasible actions and information, are both unknown, so mechanisms that make such agents act on our behalf must incentivize honesty and obedience together. A one-sided imitation structure, in which capabilities can be concealed but not counterfeited, yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. The framework is applied to stylized examples of sandbagging, where a capable agent pretends to be less capable, an alignment-interpretability trade-off in which the two are substitutes in the instrument but complements in value, discipline via peer scoring, coupled rewards that induce competition among agents, and scalable oversight with reward shaping.
4 more specialized papers
- Capability-Gated Language Models: Security Composes, Utility Does Not Patrikas Vanagas, Augustas Ma\v{c}ijauskas, Laurynas Lopata
- Causal Evidentiary Governance for High-Risk Machine Learning Systems Samah Kareem, Bar{\i}\c{s} \c{C}elikta\c{s}
- Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges Rui Yang, Shuang Huang, Junhua Liu et al.
- SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue Stephanie Fong, Yiwen Jiang, Zimu Wang et al.
Other 42
Modelpedia: A Catalog of Model Findings for the Meta-Science of AI
Findings about how foundation models behave and fail are published faster than the community can organize them and end up scattered across papers, blogs, and technical reports. Modelpedia is an automated, LLM-assisted framework that extracts findings about models from published papers, links each to the model, dataset, method, and concept it concerns, and aggregates them into a searchable public catalog. Applied to accepted ICLR 2024 and 2025 papers, the prototype extracts over a thousand findings, and the authors treat the catalog itself as an object of study to run a meta-analysis of how the community investigates models, inviting contributions to the open resource.
Superposed Latent Autoencoder
Autoencoders usually meet tight latent-memory budgets by shrinking each latent, which sacrifices representational capacity. The Superposed Latent Autoencoder (SLAE) instead keeps latents wide and shares storage: it transforms latents into storage-friendly codes, binds them with randomized keys, superposes several codes into a single memory tensor, and learns to recover each latent before decoding, replacing an irreversible dimensional bottleneck with structured interference that can be suppressed. Across CIFAR-10, CIFAR-100, SVHN, STL-10, and Tiny ImageNet, SLAE reduces reconstruction error by up to 56% over conventional autoencoders at matched storage, and the preserved information lifts downstream classification by up to 16.79 percentage points under the same memory budget.
Births are difficult to predict even with rich survey and full-population register data
Major life events are notoriously hard to predict, and it is unclear whether that reflects weak theory, data, and algorithms or the large role of chance. A data challenge had 147 researchers predict whether Dutch residents aged 18 to 45 would have a child within three years, using either rich survey data or full-population registers, with methods ranging from logistic regression to transformers and a large language model. Predictions were only moderately accurate (best F1 of 0.59 on register data and 0.76 on survey data), advanced models did not beat classical ones, and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy yields a predictive ceiling of roughly 0.86 to 0.96 F1, so observed performance falls short of what data and methods could achieve, while chance in reproduction alone still imposes a non-trivial limit on predicting individual lives.
Recent Developments in Transformer Inference Deployment on FPGA Platforms: A Survey
Deploying Transformer models for inference demands throughput, latency, and energy efficiency that Central Processing Units (CPUs) and Graphics Processing Units (GPUs) do not always deliver, and Field Programmable Gate Array (FPGA) platforms offer implementation flexibility, energy efficiency, and suitability for on-site deployment as an alternative. The authors conduct a systematic literature review of recent Transformer inference work on FPGAs, extracting the preferred implementation and optimisation techniques and organizing design choices, trends, and optimisation methods into a taxonomy. The review is intended as a guide for both academic and industry practitioners choosing how to deploy Transformers on reconfigurable hardware.
Optimizing Byzantine Node Placement in Decentralized Federated Learning
Security studies of decentralized federated learning focus on how Byzantine participants behave and largely ignore which participants get compromised, even though aggregation runs over a communication graph where placement decides how far malicious influence spreads. Treating placement as an explicit adversarial choice under a fixed compromise budget, Byzantine Placement Influence (BPI) is a set-level measure derived from the gossip dynamics that quantifies honest nodes' cumulative exposure to Byzantine sources over the training horizon, accounting for weighted multi-hop propagation and interaction among compromised nodes instead of relying on centrality heuristics. Across six heterogeneous graph families, untargeted model poisoning, and backdoor attacks, BPI-guided placements consistently identify the most damaging configurations, and remain effective when Byzantine-robust aggregation breaks the linear gossip assumption.
37 more specialized papers
- Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training Zhiyang Qiu, Yangtao Wang, Xiaocui Li et al.
- Assessing Alignment and Stability of Feature Importance Explanations via Weight of Evidence Eddie Conti, Claudio Daka, \'Alvaro Parafita et al.
- Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification Chuanhang Qiu, Yanran Xu, Yue Wang et al.
- Different representation learning objectives recover distinct latent structures from the same psychometric data Cong Cao, Tassos C. Kyriakides, Pambos Vrasidas
- Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective Behrooz Razeghi
- A Stable Aggregation Method for Quantum Federated Learning Shanika Nanayakkara, Shiva Raj Pokhrel
- Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure Filippo Cenacchi, Longbing Cao, Runze Yang
- Neural means and kernel corrections for operator learning Yitzchak Shmalo
- CRAD: Class-wise Reliability-Aware Distillation for Decentralized Heterogeneous Federated Learning Baraa Bilbeisi, Mengchen Fan, Baocheng Geng et al.
- Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach Romain Claret, Michael O'Neill, Paul Cotofrei et al.
- Context Window Failures in Relational Foundation Models Denis Oliveira Correa, Francisco Galuppo Azevedo
- Fractal dimension predicts quantum kernel collapse in angle-encoded data Ana Paula Appel
- A hybrid quantum-classical neural network for learning to route Marcus Rolf Peter Ritt, Alexsandro Santos da Rosa J\'unior, Marcos Vinicius Reballo et al.
- Wave Function Backpropagation with Explicit Temporal-Interval Dynamics Byunggu Yu, Justin Kim
- When the Algorithm Becomes the Brand Crisis: A Sociotechnical Theory of Distributed Responsibility and Accountable Transparency Mohammad Saleh Torkestani, Taha Mansouri
- A Mathematical Framework for Legacy, Governance, and Decision Integrity in Enterprise AI Shorab Sarker
- Real-Time Neuromorphic Spectrum Intelligence Simulator Navaneetha Krishnan Kamalakannan
- DK-GBMKKM: Dynamic Kernel-Space Granular-Ball Multiple Kernel $k$-Means Clustering Xiaoyu Lian, Yuchao Zhang, Shuyin Xia et al.
- A Study of Hidden-State Optimization Order in Predictive Coding Networks Xueyuan Li, Danilo Vasconcellos Vargas
- MUGEN: Generating Unlearnable Graph Examples for Multiple Learning Tasks Ziyan Liu, Chengshuai Zhao, Huan Liu
- Differentially Private Paired Table-Image Multimodal Synthesis Kai Chen, Josephine Lamp, Somesh Jha et al.
- Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures Jaee Ponde, Roshni Agarwal, Subhashis Banerjee
- When Features Become Instances: Inverted Contrastive Learning for Unsupervised Feature Selection Utsab Ghosh, Roshni Chakraborty
- Subspace Levenberg Marquardt Algorithms in Training Neural Networks M. Duc Hoang
- MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries Utsab Ghosh, Roshni Chakraborty
- FLaG: Frequency-Domain Latent-attention Gated Pooling for Token Aggregation Kewei Li, Rongying Zhang, Xueli Wang et al.
- A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling Filippo Dainelli, Amirpasha Mozaffari, Marina Casta\~no et al.
- QILP-0: Constructing Observational Declarative Twins of Quantum Circuits Marina de la Cruz Echeand\'ia, C\'esar Luis Alonso, Tony Ribeiro et al.
- Artificial Rosetta Stone: Constrained Maximum A Posteriori (MAP) Reconstruction of Symbolic Raga Sequences via Order-k Markov Models Saanvi Raghavendran (Abstract Math Institute), Abhishek Bhattacharjee (Abstract Math Institute)
- Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration Daehwan Kim, Haejun Chung, Ikbeom Jang
- Neural Symbollic Regression Using Deep Learning and Sparse Modelling Ravi Kumar U, Sumitra S
- Replicating TRACE: A Practitioner's Guide to Its Threshold and Particle Budget Alex Chadyuk, Alicia Zhang, Roy Kucukates
- Solving In-Table Prediction Problems by Deep Neural Networks with Performance Evaluation Using Synthetic Data Xiao Zhao, Daniela Oelke
- Relational Task Generation Language: A Declarative Specification Framework for Relational Deep Learning Oleksii Kolesnichenko, Jakub Pele\v{s}ka, Gustav \v{S}\'{\i}r
- Exploring Sparse Autoencoders in Text-Based Causal Confounding Adjustment Mian Zhong, Katherine A. Keith, Anjalie Field
- Contribution-Aware Bandwidth Allocation for Multimodal Split Learning Iason Ofeidis, Leandros Tassiulas
- Learning Sparse Decision Trees via Transformer Variational Auto-Encoders Giacomo Fidone, Alessio Cascione, Riccardo Guidotti
Theory 42
Dense Weak Hiding: Closing Complexity Gaps in Nonconvex and PL Finite-Sum Optimization under Individual Smoothness
For nonconvex finite-sum optimization under individual smoothness, known algorithms use O(n + √n·ΔL/ε²) incremental first-order oracle calls while the best lower bounds fell short by a factor of √n, leaving the optimal complexity open. A matching randomized lower bound closes that gap — fixing the minimax oracle complexity up to universal constants under both individual and mean-squared smoothness — and companion bounds settle the global Polyak-Łojasiewicz regime, showing restarted PAGE is optimal where the standard guarantee was not tight. The technique, dense weak hiding, spreads each hidden direction across components via a fixed sign table so an individual queried row leaks little while the exact row average preserves the signal, with a bounded radial map and a smooth gate keeping unopened links invisible to values and gradients alike.
Recursive Criticality of AI Self-Improvement
A dynamical model asks when AI systems used in AI research and development produce self-amplifying capability growth, balancing baseline research productivity, the strength of recursive feedback, and the rising difficulty of further progress. The central quantity is a recursive reproduction number whose value above one marks a regime where improvements compound across development cycles, and below one one where they are damped; the threshold does not correspond to any particular capability level, so a system can enter the amplifying regime before acceleration is visible, while rapid progress can also occur without amplification. Extending the model to multiple actors shows improvements shared between organizations can make the overall research ecosystem self-amplifying even when no individual organization is, and the framework names measurable properties — feedback strength, propagation into successor systems, cycle duration, difficulty growth — that distinguish amplification from fast progress with other causes.
Rock, Paper, Scissors, ... Dynamite - A Model of Disruption from New Technologies
To study how a disruptive new capability reshapes a competition, the authors add a versatile "Dynamite" move to Rock-Paper-Scissors and solve for equilibrium play. Giving Dynamite to only one player raises that player's win probability from 50% to just 55.5%, and it gets played rarely; the advantage shrinks further when the game is expanded beyond the original three moves. The analysis also surfaces mechanisms by which existing moves become strategically unplayable or obsolete, which the authors offer as an intuition pump for how raw capability translates — or fails to translate — into value.
Exact Global MCMC with Denoising Diffusion
Sampling from complex high-dimensional unnormalized densities is hard for local Markov chain Monte Carlo (MCMC) methods that get trapped in modes. The observation here is that running a forward then reverse diffusion process defines a Markov chain preserving the target distribution when the denoiser is ideal, and that this can be made exact for any denoiser by adding a Metropolis-Hastings correction whose acceptance ratio uses the densities of the forward and reverse paths of a discrete-time stochastic differential equation approximation. Denoising Diffusion Monte Carlo trains a standard denoising diffusion model on locally convergent Metropolis-adjusted Langevin samples and composes the resulting global path proposal with a local sampler, achieving high acceptance rates for global moves across a range of complex targets and offering preliminary evidence that diffusion training's scaling behavior carries over to exact sampling.
How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks
Prior analyses of how much memory a linear recurrent neural network builds during training assumed uncorrelated inputs, whereas real sequences are correlated. Solving the learning dynamics exactly for correlated inputs shows the entire effect of correlation collapses onto a single cost attached to retaining the past, recovering the earlier result when inputs are uncorrelated and growing under positive correlation. Three consequences follow: correlation reshapes the whole course of learning, with memory building, overshooting, and being partly removed before the network settles on retaining less; memory switches off at a threshold set by one number — how much each input resembles the one immediately before it — independent of sequence length or longer-range correlation; and given a single spare hidden dimension, training spontaneously builds a feedthrough path that passes the current input straight to the output and remembers nothing. The analysis turns one property of the input into a prediction of whether memory is learned at all, and explains why correlated data turns recurrent networks into change detectors.
Independent Reinforcement Learning in Discounted Markov Games
The setting is radically uncoupled learning in discounted general-sum Markov games, where each player updates from its own rewards with no knowledge of the others. Assuming the exponential time hypothesis for the complexity class PPAD, the authors prove that for every fixed discount factor no polynomial-time independent-learning algorithm can compute inverse-polynomially accurate coarse correlated equilibria. They then give what appears to be the first radically uncoupled algorithm with sub-exponential convergence to coarse correlated equilibria in this class of games without structural assumptions: a layered variant of optimistic mirror descent with an increasing step-size schedule, in both full-feedback and partial-feedback versions.
Disciplined Bilevel Programming
Bilevel optimization models hierarchical decisions where one problem is nested inside another, but using existing solvers demands manual reformulation by an expert. Disciplined bilevel programming (DBLP) lets users write optimistic bilevel problems close to their mathematical form, then automatically canonicalizes a convex lower problem into conic form and builds an equivalent single-level reformulation via conic Karush-Kuhn-Tucker conditions, relaxing the complementarity constraint and solving a sequence of smooth nonlinear problems through gap continuation. The implementation, BLVPY, extends CVXPY and lets users state and solve bilevel problems in a few lines of code without bilevel modeling expertise, demonstrated across several application domains.
Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets
A service fronting several language models must choose one per request while respecting workload-level budgets for compute, latency, memory, or cost, and both the request mix and the model lineup drift after launches, fine-tunes, and quantization changes. The setup is formulated as nonstationary sparse contextual routing under multiple knapsack constraints, with an optional shadow-audit stream that scores a small fraction of prompts on several models. DRS (Drift-Aware Sparse Routing) estimates reward and resource use from a rolling audit window, routes on pessimistic reward and optimistic cost, updates resource shadow prices online, and applies a hard budget meter before committing. Regret against a paced dynamic benchmark is bounded in terms of sparsity, audit rate, window length, and drift, recovering the standard O(√(sT/ρ)) rate when there is no drift and a T^(2/3) adaptation term when there is.
Verdict Instability of OOD Scores under Reference Resampling
Post-hoc out-of-distribution detectors are fitted on a finite reference set, so every score is an estimate and a different reference sample would move some verdicts; the authors measure that movement as the bootstrap standard deviation of the score, called verdict instability, and derive a closed form with no fitted parameters. Instability turns out to be the within-class dispersion of the assigned class along the query's direction divided by the square root of that class's reference count, and the count term is identifiable only under class imbalance. Because far-OOD queries lie along the low-variance directions of anisotropic embeddings, every distance-based score tested assigns its most extreme values to its most reproducible verdicts, and only estimators of local dispersion carry the sign practitioners expect. A single label-free correlation predicts that sign for any score, and abstention driven by a wrong-signed score is worse than abstaining at random on every dataset tested.
Denoising Diffusion Generative Models Secretly Calculate Attentions
Denoising diffusion models dominate image generation while attention-based transformers dominate language modeling, and the authors argue the two families share more than they appear to. They show that diffusion models inherently compute an attention-like operation close to the transformer's, and draw a parallel between auto-encoders and attention-based models, suggesting the designs can be swapped depending on practical requirements. Using this equivalence, they reformulate the diffusion framework into a simplified attention-based image generation algorithm and report that it achieves comparable performance with significantly less training effort and compute.
When Metropolis and Hastings Meet Bradley and Terry: Exact MCMC From Preference Voting
Sampling from a generative model's distribution conditioned on a semantic property is a natural fit for Metropolis-Hastings (MH), but MH needs exact pointwise density ratios, which are unavailable in generative settings where only pairwise comparisons from humans or model judges are cheap. Pref-MH observes that the MH unnormalized density ratio equals the preference odds under the Bradley-Terry (BT) choice model, and develops an accept/reject rule that works from sampled binary judge feedback while still provably converging to the target distribution. The authors show that for a fixed proposal kernel and comparison budget, the rule is Peskun-Tierney optimal among exact reversible acceptance rules of this class. Experiments on text generation and molecular design with LLM judges, and on image generation with vision-language model judges, demonstrate practical conditional sampling from comparative feedback alone.
The Multiple Timescales of Gradient Descent on the Edge of Stability: A Perturbative Derivation of the Central Flow
The central flow of Cohen et al. is an empirically accurate continuous-time model of gradient descent at the edge of stability, but its original derivation was heuristic. The authors propose a perturbative regime in which the loss decomposes as a main term plus a small perturbation, treat gradient descent as a singularly perturbed dynamical system, and apply the classical method of multiple scales to expand the dynamics. Three timescales emerge, fast oscillations along the sharpest direction, an intermediate self-stabilization mechanism, and slow motion along the minimizers of the main term, and the central flow appears as the leading-order term of the expansion with self-stabilization at the next order; the analysis also computes the slow drift of fluctuation energy and explains why fluctuations persist when several eigenvalues sit at the edge of stability.
Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras
A sparse subset of Transformer attention heads has effective output-value (OV) operators that nearly close under composition, with the operator squared approximately equal to a scalar multiple of itself, a property the authors call scaled idempotence. Across six pretrained models from 2.8B to 235B parameters, 3.98 to 8.00% of heads reach a squared-closure alignment of at least 0.9, while no matched within-layer O/V mismatch does, and an exact principal-coordinate factorization separates within-support transport from read-write return geometry. Scrambling only the orientation of the core factor while preserving singular values, norms, factor spans, and principal angles collapses median closure from 0.336 to about 1e-4 across 7,304 heads in nine multi-head and grouped-query attention models, showing the property is a trained orientation within broadly available geometric capacity. Under exact value sharing, headwise closure extends to a right-action algebra across heads, verified approximately in seven models.
On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study
Generative augmentation is a standard remedy for class imbalance, but its effect on downstream generalization lacks a theoretical account. Framing augmentation as a distribution-mixing process shows the resulting risk distortion is controlled by augmentation strength and the class-conditional Wasserstein discrepancy between real and generated data, and a Rademacher-complexity bound makes explicit the trade-off between hypothesis capacity, augmentation intensity, and generative fidelity. Experiments with Conditional GAN and Conditional WGAN-GP on binary and multiclass imbalanced tasks confirm that CWGAN-GP achieves lower Wasserstein discrepancy, but higher generative fidelity does not reliably translate into better classification, with classical oversampling often remaining competitive.
28 more specialized papers
- Convergence issues in Relational Concept Analysis based on AOC-posets Xavier Dolques, Agn\`es Braud, Alain Gutierrez et al.
- When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation Cong Cao
- Stochastic complexity of vectors containing cluster structure Daniel Nicorici, Olli Yli-Harja, Jaakko Astola
- Flawed in Nature, Perfect through Evolution J. M. Diederik Kruijssen (Allora Foundation)
- Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost Zihang Liang, Haochen Zhang, Lingzhou Xue
- Towards unsupervised representation learning for quantum data: quantum models with inference and generation Robin Lorenz, Eric Brunner, Marcello Benedetti
- Operational Regimes in Non-Convex Optimization: A Multiplier-Based Taxonomy Seyed Mohsen Kazemi, Ali Movaghar, Shaahin hessabi
- Higher Structures in Deep Learning Michael L. Roberts, Carlos Zapata Carratal\'a. Nicholas J. Cooper, Lijun Chen et al.
- Why Multi-Layer Message Passing Works: Completeness Theory for Graph Neural Network Interatomic Potentials Pingbing Ming, Han Wang
- Manifold-Aware General Coded Computing for Straggler-Resilient Distributed Computing Parsa Moradi, Mohammad Ali Maddah-Ali
- Prediction-Assisted Pricing and Admission for LLM APIs with Stochastic Token Consumption Patrick Wong
- Measuring Optimal Transport in Transformer Depth Alexandre Quemy
- Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan
- Sharp Mixed Spectral Barron Regularity of Coulombic Many-Electron Wave Functions Pingbing Ming, Hao Yu
- Poisson-Gamma Dynamical Systems with Time-varying Transition Dynamics Jiahao Wang, Yijun Wang, Nan Fang et al.
- Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches Marco Simnacher, Georg Keilbar, Benjamin K\"onig et al.
- On Synthesis of Metric Interval Temporal Logics Hsi-Ming Ho, Shankaranarayanan Krishna, Khushraj Madnani
- Multi-Head Self Attention is a Parameter Identification Mechanism W. Ross Morrow
- One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context Skanda Athreya, Yutong Wang
- Matched Queries for Curvature and Density at Branching Junctions Ziqi Zhao, Qingjian Ni
- Exact Risk-Complexity Laws for Projective Boundaries in Scenario Optimization and Distribution-Free Certification Giuseppe C. Calafiore
- Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate
- Edge-Girth as a Structural Edge Feature for Graph Neural Networks Lilian Marey, Charlotte Laclau
- Rethinking Learnability in Offline Data-driven Optimization Chao Qian, Chen-Guang Wang, Rong-Xi Tan et al.
- Sierpi\'nski--Knopp Wasserstein Distance for Persistence Diagrams and Applications to 2-Wasserstein Approximation Sebastien Tchitchek, Julien Tierny
- Variable Selection for Feature-Based Newsvendor Zhaoliang Yuan, Jie Wang
- A Mathematical Theory of Reusable Neural Bases for Network Compression Binshuai Wang
- Gradient-Update Mismatch: Rethinking Conflict-Free Training of Physics-Informed Neural Networks Jing Xiao, Xinhai Chen, Qinglin Wang et al.
Multimodal 40
SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces
Architecture drawings, flowcharts, and pipeline schematics often carry more information than the surrounding prose, yet no public corpus pairs that figure type with the captions, context, questions, answers, and reasoning steps needed to train a vision-language model on them. SCAFFOLD builds such tuples from arXiv computer science papers using layout detection and PDF parsing plus an AI-assisted question-generation stage, yielding 157,387 pairs across 29,887 figures from 3,058 papers in the largest split, with 37K and 12K subsets alongside it. Baseline experiments use the 12K subset on Qwen2.5-VL-3B-Instruct.
Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
Multimodal LLMs will let conflicting text override what an image plainly shows, a failure the authors term multimodal contextual sycophancy and probe with a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text. The key manipulation is where the information boundary sits around a context-blind visual witness: on abnormal images paired with false Gemini-generated text, GPT-5.1 scores 7.9% when conditioned jointly, 49.7% when the context-blind witness report is scored directly, 63.7% in a two-call witness-arbiter pipeline that shows the witness the text, and 84.2% under System-2 Visual Arbitration, which withholds the text from the witness. Across six models that last setup beats the direct witness report by 19.7 to 44.1 points with all paired confidence intervals excluding zero, though the best boundary is model- and source-dependent — text helps some models, and a GPT-4o-regenerated subset reorders the conditions.
Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy
Memory benchmarks for language and vision-language models (VLMs) usually report accuracy over long text or video, which reveals nothing about how much computation an answer costs or whether the system knows when to abstain. ECCBench adds three axes to capacity: efficiency measured in FLOPs needed to answer from memory, compression (whether compressible inputs are remembered more accurately or more cheaply), and calibration (abstaining in proportion to uncertainty and the cost of an error). Pretrained VLMs turn out to compress their memory over text but not over video, and are poorly calibrated on both, while several non-Transformer memory backbones achieve better compression-calibration tradeoffs than RoPE Transformers, suggesting them as components for long-horizon agents.
Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
Multimodal large language models used for video moderation miss a compositional failure: a video whose individual parts are all benign can still convey harmful meaning as a whole, a phenomenon the authors call Distributed Implicit Harm and study along two axes — harm distributed across temporal visual segments and harm arising between audio and visual streams. Because such videos are absent from safety datasets and resist retrieval by keywords or local visual cues, the authors built a multi-agent synthesis pipeline that composes benign components into harmful scenarios, producing over 9,000 videos with explicit reasoning annotations. Benchmarking more than 30 models, including frontier proprietary systems, revealed consistent deficits on both axes: models often judge each component correctly in isolation yet fail to see the meaning emerging from their combination, and the same failure appeared on real social-media videos collected manually.
CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction
Long-tail driving failures are usually treated as rare-object recognition problems, but the decision-relevant question is how an unusual object constrains the ego vehicle's feasible actions. The authors define decision-level driving affordance prediction, mapping a front-view image, ego-motion history, and navigation command to a structured longitudinal-lateral meta-action, and release CoLT-Drive, a 3,536-sample counterfactual benchmark that inserts rare objects into otherwise fixed scenes. Their adaptation method KPA combines perception-to-decision prompting, spherical-interpolation expert merging, and RegMoE, a regime-aware mixture of low-rank adapters, to add task capacity without erasing pretrained open-world knowledge. On CoLT-Drive it reaches 60.8% action-pair accuracy versus 50.3% for the Qwen3-VL-2B baseline and 32.4% for plain LoRA supervised fine-tuning, which degrades sharply.
Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts
Vision-language models must sometimes choose between information supplied in context and facts memorized during training, and the resolution turns out to depend on modality. Models tend to follow in-context information for entities named in text but fall back on memorized facts for entities shown in images. The authors attribute this to late representational alignment across modalities: resolving a visual entity takes longer, so the model's usual factual-recall mechanism is never suppressed in time, yielding parametric answers. Chain-of-thought prompting does not close the asymmetry, though adding more visual information to the context does shift behavior.
Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models
Speculative decoding speeds up generation without changing outputs, but on vision-language models it has been stuck in a loop: the drafter is autoregressive so it must stay small, a small drafter cannot process the image every step, and a vision-starved drafter fails exactly where the image would make text predictable. GLANCE breaks this with a block-diffusion drafting head that reads the target model's already-fused vision-language state, so images cost the drafter nothing, and fills an entire block in a single forward pass, with a wide candidate tree verified in one target pass and every audited prompt reproducing greedy decoding exactly. Under one engine and round budget it decodes up to 2.93x faster than autoregression using one draft pass per round where the production EAGLE3-VL head takes eight, and accepts blocks 2.7x longer than an EAGLE-3 head trained on the same data. The authors also fit a relationship between accepted length and the target's next-token entropy whose slope steepens with grounding, holding across five tasks and multiple targets while predicting where free-running text still favors chain drafting.
(V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement
Whether language models learn abstract grammatical rules or merely track lexical co-occurrence is hard to settle with text alone, since distributional cues like is/are and this/these give grammatical number away. The test here uses vision-language models (VLMs): new nouns are taught by adding fresh embeddings and updating only those, comparing a condition where singular versus plural is signalled purely by the image against one where text disambiguates it. Across behavior, representational dynamics, and causal interventions, the models show non-trivial cross-modal generalization in both exposure conditions, and the internal mechanisms treat linguistic and extra-linguistic cues in similar ways, which the authors read as abstraction-compatible behavior rather than surface pattern matching.
Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict
The same evidence can be handed to a multimodal model as text, as a rendered image of that text, or as both, and it is unclear whether these surface forms are treated equivalently when the evidence contradicts what the model already believes. Testing 13 multimodal large language models on two datasets under knowledge conflict, the authors find that models more readily accept contradicting context in image form than in text form, and that when conflicting text and image appear together the winning modality is essentially arbitrary, shifting with input order, model, and dataset. The instability degrades multimodal retrieval-augmented generation and is exploitable by adversarial attacks; of the mitigations tried — prompting, activation steering, supervised fine-tuning, and direct preference optimization — only supervised fine-tuning helps moderately.
EM^2Mem: Event-Centric Multimodal Memory for Large Language Models
Memory systems for long-video question answering usually store captions, frames, transcripts, and graph facts as separate fragments, forcing the language model to reassemble cross-modal and temporal alignments at inference time when context is scarce. EM^2Mem instead binds heterogeneous evidence to event anchors while memory is being built, so each event-indexed cell already carries aligned multimodal records, temporal context, graph relations, semantic facts, and provenance. Across three long-video QA benchmarks it gains 2.0, 2.4, and 3.7 accuracy points over the strongest memory baseline and 7.0 points of strict event-level top-5 evidence recall, while cutting per-query latency 4.67 times and total inference tokens by 63.66%.
VoiceLongMemEval: Do Assistants Remember How You Sounded?
Long-horizon memory benchmarks for assistants test what was said across sessions but not how it was said, ignoring emotion, prosody, and voice events. VoiceLongMemEval makes every answer depend on paralinguistic metadata attached to conversational turns, with a three-stage adversarial gate ensuring a strong text-only language model fails on each item. Evaluating frontier and open-weight models exposes a pervasive affect gap: supplying the paralinguistic metadata as text raises accuracy by 0.09 to 0.38 (reaching 0.61 to 0.69 with evidence hints), audio-native models recover some of the signal directly from speech (0.354 to 0.412 versus 0.325 blind), and standard speech-recognition pipelines discard it entirely.
DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
Commercial short-drama production runs through script, storyboard, keyframe images, shot-level video, and final assembly, yet existing benchmarks score only the video-generation stage using hand-authored inputs rather than actual upstream outputs. DramaChain Bench evaluates every stage with a shared set of five axes expanding into 63 leaf dimensions, backed by 5,785 items each scored by three professional annotators, producing 17,488 scores and 255,925 attribution records with defects localized in space and time. The human annotations confirm that upstream defects cascade downstream, so final episode quality is not determined by video generation alone, and an agentic automatic judge that gathers evidence over multiple rounds reproduces the human model ranking at a mean PLCC of 0.918.
SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task
The NTCIR-19 SciClaimEval task asks systems to verify scientific claims against the tables and figures of a paper, and the entry described here benchmarks eleven frontier and open multimodal models under one per-sample protocol instead of tuning a single system, then adds light post-processing. The team placed first in three of four evidence-category and subtask combinations, with GPT-5.5 and Claude Fable 5 leading both subtasks and Claude Opus 4.8 and Gemma-4-31B beating the strongest public baseline. The largest gain came from the data's structure rather than model choice: a leak-free pair prior raised Subtask-1 pair accuracy from 72.2 to 93.5, more than any model swap or ensemble. A case-by-case audit attributes most remaining errors to label noise or label-mapping swaps, and documents a measurement leak in which released file ordering encodes the answer.
Controllable Image Captioning with Prompt-Conditioned Scene Rewards
Large vision-language models write fluent image descriptions but give users little control over whether a caption emphasizes attributes, relations, or particular regions. FoCUS adds a prompt-conditioned reward: generated captions are parsed and aligned to scene-graph components such as objects, attributes, and relations, those components are weighted — negatively when the user asks to avoid something — according to the control prompt, and the model is optimized against the resulting objective with GRPO, backed by a stricter object validity threshold and reasoning-based verification for attribute and relation scoring. The authors also introduce SCoPE, a benchmark of contrastive Include/Avoid constraints that scores both coverage of requested content and suppression of out-of-scope content, and report improved controllability and fine-grained caption quality on two VLM backbones without degrading general captioning.
Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
Audio language models are meant to understand speech, but it is unclear whether they represent how something is said rather than only what is said. Using the Expresso dataset of controlled speaking styles, the authors trace paralinguistic information through Whisper-large-v2, Qwen2-Audio-7B-Instruct, Qwen2.5-Omni-7B, and Chroma-4B with centered kernel alignment, leave-one-speaker-out linear probing, open-ended tone prediction, and a content-prosody leakage metric. All four models encode speaking style strongly in the top third of the audio encoder, yet that information is consistently degraded before it reaches the output: the projector reshapes geometry without destroying style, while decoders split into content-driven behaviour, where predictions track the text, and acoustic-driven behaviour, where they vary with delivery.
Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis
Turning a pretrained language model into a vision-language model erodes its text ability, with damage concentrated on tasks graded by strict output formats such as instruction following and strictly parsed chain-of-thought answers. The authors attribute this to attention-sink corruption: visual fine-tuning perturbs the early token position that absorbs a large share of attention probability, and how well a backbone preserves that sink predicts how much capability survives. They introduce Sink Strength, a single scalar computed on the base model in seconds on one GPU that tracks relative post-adaptation degradation across six model pairs without any vision-language training, and report two negative results: injecting QK-RMSNorm after pretraining does not reproduce the protection of native QK-RMSNorm, and several off-the-shelf weight-merging recipes fail to recover the lost ability.
Solaris: Towards Interfaces That Are Generated, Not Coded
User interfaces are normally specified ahead of time through code, which fixes their appearance and behaviour to a predefined set of states. Solaris is an interface world model that skips that intermediate representation entirely: it treats mouse input as a conditioning signal and autoregressively generates each frame of the interface, with few-step distillation and training on its own outputs used to hold visual coherence at interactive speeds over long sessions. A separate language model interprets user intent and specifies how interactions should change the environment, splitting high-level reasoning from visual rendering so that interactions which were never programmed can still be carried out.
Visual Attention Faithfulness in Vision-Language Models is Heterogeneous
Whether attention weights explain a model's reasoning has been argued over for years in natural language processing, but the question is largely untested for the visual side of vision-language models (VLMs). Causal perturbation analysis measuring both comprehensiveness and the sufficiency gap of attention-ranked visual tokens finds that faithfulness is not a single property of a model but splits into three processing modes: Faithful-Sufficient, where the top-k attended tokens are both necessary and sufficient; Faithful-Distributed, where they are necessary but wider context is still needed; and Non-Focal, where no localised region is individually necessary even though visual input remains essential. Human-annotated ground-truth regions satisfy comprehensiveness in only about 60% of cases relative to model attention rankings, and the mode a model falls into varies systematically with architecture and task across VQAv2, VRDU, and ChartQA.
Towards Generalizable Visually Grounded Exploration of Household Devices
Operating an unfamiliar household appliance without a manual requires forming hypotheses, acting, and correcting from feedback, a loop that existing embodied benchmarks sidestep by supplying documents and annotated demonstration trajectories. VGEBench targets what the authors call generalizable visually grounded exploration by building a logic-driven state machine that simulates multi-turn interaction, forcing a vision-language model to reach goals through active perception and feedback-driven correction rather than imitation. Evaluation shows current vision-language models struggle to translate semantic world knowledge into correct physical action sequences and to track device state over long horizons.
Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation
Generating a pathology report from a whole-slide image is held back by the scarcity of paired slide-report data and by the difficulty of turning spatially scattered visual patterns into structured clinical text. A clinically curated Pan-Asia dataset of roughly 10,500 slide-report pairs from five institutions underpins the REG 2025 benchmark, run as a MICCAI challenge, whose submissions span pretrained vision-language models, multiple-instance learning, hierarchical expert models, retrieval-augmented generation, and cross-modal transformers. The analysis finds that using a vision-language model was not itself what separated top entries — structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding were — and identifies recurring failures including numeric hallucination in attribute estimation and diagnostic overspecification that mirrors known pitfalls in routine practice.
The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence
Reporting aggregate accuracy on a multimodal benchmark quietly assumes the model actually looked at the image, and that assumption is tested by blurring the question-relevant region and measuring how much the next-token distribution moves. Across six vision-language models and three perceptual benchmarks, the distribution barely shifts on 40% to 97% of samples, a phenomenon named the Visual Insensitivity Gap and scored per sample by a Visual Sensitivity Index; the effect belongs to samples rather than models, since the index correlates across architectures sharing nothing but a contrastively pretrained vision tower (grand-mean Spearman rho = +0.40). A linear probe on each model's own vision tower separates perturbed from clean images at 0.72-0.79 accuracy on exactly those samples while the model's top token changes on only 2-11% of them, locating the failure in the encoder-to-language-model handoff rather than in perception, and the index works best as a conditional signal for multiple-choice reasoning on capable models (AUROC 0.85-0.87) rather than as a universal abstention criterion.
From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding
Vision-language models (VLMs) handle natural-image question answering well but struggle with scientific diagrams, which convey functional or relational meaning rather than literal scenes. The authors propose a framework that extracts domain concepts from science-curriculum terminology, synthesizes atomic facts, retrieves relevant diagrams from the web, and generates captions and multiple-choice questions as multimodal supervision. The resulting SciGram dataset holds over 194K diagrams and 1.4M visual instructions across life, earth, and physical sciences, and despite noisy web data and synthetic labels, models fine-tuned on it match or outperform state-of-the-art VLMs while using fewer training instances on TQA, ScienceQA, and AI2D. Augmenting LLaVA OneVision with SciGram sets new state-of-the-art results on diagram question answering, and both the dataset and the models are released.
SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models
Multimodal large language models pay a heavy compute cost for long visual token sequences, and existing pruning methods often mistakenly preserve high-norm outlier tokens that are actually redundant in both feature and spatial dimensions. SinkPruner is a training-free coarse-to-fine framework with a visual sanitizer that filters these high-norm redundancies while easing attention sink and dispersion, followed by a text-guided pruner that keeps tokens semantically aligned with the query. Across twelve image-language and four video-language benchmarks, it preserves 96.5% of LLaVA-1.5 performance and 91.8% of Qwen2.5-VL performance under an 89% token reduction, and the sanitizer also improves existing pruning methods when transplanted.
When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP
Shrinking the modality gap between image and text embeddings in CLIP is widely expected to improve zero-shot accuracy, yet a smaller average gap often fails to deliver consistent gains. The authors analyze this through the decision structure of zero-shot classification, where accuracy depends on class-wise decision margins rather than average alignment, and use linear correction as a tractable case to show that gap correction can reshuffle relative margins so that predictions collapse onto a small subset of classes, a failure mode they call prediction-level hubness. Experiments across multiple datasets show that accuracy drops under gap correction are consistently accompanied by increased prediction concentration, for both linear and learning-based correction methods, so the authors argue gap correction should be judged by its effect on downstream prediction structure rather than by average alignment alone.
On the Design Fundamentals of Pixel Text Representation Learning
Text-rich visual inputs require encoders that read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders suffer from fixed-resolution pretraining, visual shortcut learning, weak grounding, and poor multilingual coverage. Controlled ablations identify four design principles: variable image resolutions and font sizes as proxies for high-resolution documents, natural image-text pairs to prevent text-only collapse, layout-aware rendering to block pixel-level shortcuts, and a two-stage multilingual curriculum for cross-lingual alignment. These are combined into Pixel Linguist II, a native-resolution vision encoder trained with on-the-fly rendering and unified contrastive grounding over 280M examples, which sets new state-of-the-art results on English, cross-lingual, and multilingual Visual STS and ViDoRe and improves downstream multimodal LLM evaluation. The encoder remains robust under 80% visual token compression, pointing to optical context compression as a use case.
A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation
Evaluating an omni-modal foundation model across text, image, video, and audio currently means juggling separate toolkits whose inference engines, prompt conventions, and metric implementations are mutually incompatible. OmniEvaluator connects existing inference engines and curated evaluation libraries at a higher level rather than reimplementing benchmarks, exposing four inference backends, four evaluation frameworks, and over a thousand benchmarks through a single interface, recording every run as a fully reproducible configuration artifact, and feeding results into a shared dashboard for cross-model comparison. A federated mode shares GPU inference servers across concurrent evaluations, and a built-in verifier small enough to run on CPU keeps scores stable across engines and prompts where rule-based scoring fluctuates, matching cost-efficient commercial LLM judges without recurring API cost.
MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval
Visually rich documents hide key content in tables, charts, figures, and layout that plain OCR corrupts or drops, while ColPali-style visual retrievers fix this with patch-level multi-vector indexes and late-interaction scoring that keep image-derived retrieval on the query-time serving path. MIDR (Multimodal Indexing for Document Retrieval) is a training-free framework that shifts multimodal reasoning to index time, using a multimodal LLM at ingestion to convert rendered pages into verified textual fields indexed with BM25F and optionally fused with dense retrieval. On ViDoRe V3, MIDR Hybrid reaches 0.6219 average nDCG across five English domains, a 23.0% relative gain over BM25, and enrichment lifts BM25 on French documents with English queries from 0.1532 to 0.5448. Across all seven domains it leads ColQwen2.5 on four while using roughly 9x less index memory and 2x lower query latency.
Reliability Challenges in Diffusion Vision-Language Models
Diffusion-based large vision-language models (dLVLMs) offer parallel decoding, bidirectional context, and controllable generation as an alternative to autoregressive (AR) models, but their reliability has not been characterized. The authors benchmark six diffusion models against competitive AR baselines on hallucination and bias across four dimensions. They find that dLVLMs reverse the yes-bias of AR models on binary visual questions, match AR hallucination rates but with degraded linguistic quality, collapse to near-zero accuracy on underrepresented racial groups with opposite-polarity gender bias, and lose accuracy on multiple-choice questions whenever the correct option is shorter than its distractors, a length prior that appears at the first denoising step. Tokens committed late in denoising with low confidence correlate with hallucinated content, a mechanistic signal specific to diffusion generation, and the patterns vary across model families.
EdiTikZ: Scientific Figure Editing from Revision Trajectories
Publication-ready scientific figures come from iterative refinement, yet figure editing with vision-language models (VLMs) is largely unexplored, and existing training supervision relies on costly proprietary agent systems or synthetically generated edits. DaEdiTikZ instead mines naturally occurring revision and development trajectories, collecting 391K plausible TikZ edit pairs from arXiv, GitHub, and TeX Stack Exchange and having a VLM infer 781K directed edit instructions from rendered figures and code, paired with a 790-instance human-refined benchmark. Two compact Qwen3.5-based EdiTikZ models (4B and 9B) trained on joint reconstruction and editing followed by reinforcement learning with rendered-fidelity and edit-application rewards put the 9B model above all tested baselines automatically and, across 4,320 human ratings, above GPT-5.6-Sol and on par with Gemini-3.1-Pro.
TempCloze: Can Video-LLMs Identify the Missing Middle?
Temporal reasoning benchmarks for video language models are often mediated by language, letting models exploit option wording, answer correlations, or language priors instead of watching. TempCloze reframes the task as a cloze test: given a video's beginning and ending clips, pick the true missing middle from four same-source distractors constructed along Semantic (what should happen), Alignment (when it should occur), and Progression (how it unfolds) axes, with shared scenes and objects suppressing appearance shortcuts, across 1,521 filtered mostly long-take and egocentric videos. Testing 10 proprietary and 21 open-source models identifies Alignment as the primary bottleneck: they recognize plausible content and local event progression but cannot place events correctly in time. Additional TempCloze-Mixed and TempCloze-Hard splits analyze error patterns and sensitivity to candidate order, context direction, visible span, frame density, and test-time scaling.
H3-World: Turning Language Understanding into World Control
H3-World turns the 33B MiniMax-H3 video generator into an interactive world model by exploiting the observation that large video generators already accept coarse natural-language control of character behavior and camera motion. Each action is represented as a structured combination of character and camera instructions aligned to the corresponding temporal video latents, and temporal attention routing confines each instruction to its intended time interval to reduce control leakage across actions, with no dedicated action modules added. Using only 8,000 gameplay samples, 10,000 LoRA optimization steps, and 0.199% trainable parameters, the system achieves effective character and camera control, preserves generation quality, and generalizes to unseen scenarios.
Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics
Extracting structured fields from hundreds of millions of documents a year stays expensive in regulated industries, where bespoke OCR cascades cover few workflows, privacy rules bar external models, and open-source vision-language models (VLMs) that meet quality thresholds cost more to serve than human annotation. The deployed system fine-tunes a Mixture-of-Experts VLM with 35B total and 3B active parameters on in-house production data mixed with open-domain documents selected by a difficulty-aware curation pipeline targeting layout diversity, fact extractability, and cross-model consistency, and it serves heterogeneous workflows through prompting on a single H100. It leads all deployable non-reasoning baselines up to an order of magnitude larger, and a quality-adjusted cost analysis calibrated from production telemetry shows expected costs falling by over 80% against the human baseline and by more than 50% against the best competing open-source model, while larger baselines remain economically unviable.
8 more specialized papers
- SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning Beidi Zhao, Gexin Huang, Ciro Zhang et al.
- MemeBridge: A Dataset for Benchmarking and Mitigating the Bidirectional Cultural Gap in Meme Interpretation Hangxiao Zhu, Suliu Qin, Zhuoyan Li et al.
- ExpArt-KG: Artwork Image Description Generation through Iterative Exploration of Knowledge Graphs Yuta Kato, Shintaro Ozaki, Kazuki Hayashi et al.
- Joint Training Is Not Enough: Conditioned Cross-Granularity Training for Multimodal Document Understanding Chengguang Gan, Yunhao Liang, Hanjun Wei et al.
- TEIDAN: A Multilingual Multiparty Dialogue Corpus Taiga Mori, Koji Inoue, Mikey Elmers et al.
- Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning Yuanjun Zhang, Fuzel Ahamed Shaik, Suvojit Acharjee et al.
- TimeSteer: Inference-Time Speech Scheduling in Joint Audio-Visual Diffusion Models Chao Zhou, Yiling Chen, Qi Chu et al.
- AutoConcept: Training-Free Concept-Guided Reranking for Metadata-Available Composed Image Retrieval Tianyu Wang, Tianjiao Wu
Vision 17
ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration
Reward post-training of diffusion image generators piles probability mass onto a few reward-favored modes, destroying within-prompt diversity, and existing fixes rely on external signals such as perceptual objectives or text-encoder changes rather than repairing an adapter that has already collapsed. Starting from the observation that online post-training reallocates mass over pretrained capabilities instead of learning new content — so collapse is suppression rather than deletion — ReNFT recalibrates from inside the generator: unconditional probes pick anti-hub prompts where prompt-independent bias is most visible, two mixed rollout routes produce matched counterfactuals from the same prompt and noise, and reward ranking with an adaptive flipping guard assigns pull and push roles for a paired update. On PickScore and GenEval it keeps 98.9% and 99.0% of the reward achieved by NFT while raising DreamSim diversity by 58.8% and 55.0%.
Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You
Test-time adaptation normally assumes weights can be updated at inference, which rules out inference-only accelerators, frozen third-party models, and architectures without BatchNorm where standard configurations go inactive. CASTER keeps the network frozen: it stores source class statistics in a discriminative subspace, estimates a class-shared affine transformation from target-batch moments, and analytically transports the source class distributions before classification, with no backward pass, optimizer state, or stored source feature bank. It beats k-nearest-neighbours on identical frozen features in 27 of 28 backbone-dataset settings while retaining a median of 18x less state, but transport is not always safe — on ImageNet-C, where 64-sample batches must cover 1000 classes, unconditional transport loses 21.2 top-1 points. Gating on an empirical residual-to-margin transportability certificate converts an average -3.35-point effect into a +1.69-point gain, though the authors note it does not cleanly separate benign from destructive regimes and is specific to this mechanism rather than transferable to methods like Tent.
Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation
Pixel-level segmentation of aerial and satellite imagery for tasks like flood mapping or damage assessment suffers when general foundation models miss classes or small objects in a scene. The proposed pipeline runs a vision-language model (VLM) on a single consumer-grade GPU to supply guidance without retraining: a frozen foundation model labels every pixel, one VLM query selects which classes are relevant to the scene, and a second locates small objects the base model overlooks. Evaluation on four aerial datasets shows consistent gains at each stage where the base model is already competent, plus structured evidence that can be audited independently of the mask.
VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM
Online systems that build open-vocabulary 3D maps typically segment and name an object the moment it is first detected, committing to a decision when visual evidence is weakest. VOIM (Voxel-Grounded Online Instance Manager) instead accumulates soft evidence from off-the-shelf perception models per voxel across views and defers both the instance grouping and the label until that evidence has settled, requiring no training and running from RGB-D or from monocular RGB alone. On ScanNet++ under a matched protocol it reaches 44.07 mIoU versus 32.37 for the strongest online RGB-D baseline, OVO-SLAM, winning all ten scenes, and the same system transfers unchanged to fully monocular input, matching that baseline on Replica. Swapping the region descriptor, detector label prior, and mask source shows the mapping stage rather than the perception models drives the gain, though labelling itself is not real-time because per-class detection over the full vocabulary dominates the cost.
Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking
Monocular video flattens 3D scenes into 2D projections, and multi-object trackers that rely on image-plane appearance and geometry inherit the resulting depth and spatial ambiguities. PLANET is an end-to-end tracker that first lifts existing 2D tracking datasets into 3D, then forms world-grounded queries by embedding reconstructed scene geometry into the features and positional encodings used to build queries. An auxiliary 3D location prediction task pushes queries to encode object positions during training, and a dual-resolution temporal memory preserves that evidence across longer gaps, yielding state-of-the-art results on three diverse tracking benchmarks.
ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives
Contrastive self-supervised learning has been overshadowed by generative and self-distillation approaches for pretraining vision transformers. ViTAMINS integrates synthetic hard negatives into unsupervised vision transformer pretraining through simple modifications to existing contrastive frameworks, and is benchmarked on ImageNet, transfer learning, image retrieval, copy detection, and image and video segmentation. The synthetic negatives give rise to representations that explicitly encode semantic content and serve as strong classifiers with gains of up to 11.3% over baselines, and a ViT-B trained this way surpasses V-JEPA with the larger ViT-L while using fewer resources.
11 more specialized papers
- Soft-Argmax for the Projective Plane via the Veronese Embedding Benjamin El-Zein, Dominik Eckert, Paul Zech et al.
- ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection Tongtong Wang, Mingzhu Xu, Chenglong Yu et al.
- Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting Udo Schlegel, Shubhangi, Gabriel Dax et al.
- Semi-Supervised Virtual Staining via Morphology Preservation and Histopathological Realism Constraints Baoshun Wang, Weiping Lin, Linwu Wang et al.
- SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations Yiming Luo, Rongqiang Zhao, Jie Liu
- Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test Set Michael Zang, Haiyu Wu, Mrinal Sharma et al.
- StainPresetNet: Stain Preset Network for Fast Multi-to-Multi Stain Normalization Hongtao Kang, Die Luo, Li Chen et al.
- Position Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling Stefano Leggio, Giulio Rossolini, Alessandro Biondi
- HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention Primitives Sathiyamohan Nishankar, Pubudu Sanjeewani, Asanka Perera et al.
- Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations Qingde Li, Qingqi Hong, Jie Tian
- BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Conditioned ResEnc U-Net Marven Sherif (Brightskies), Amgad Elmasry (Brightskies), Youssef Ghazal (Brightskies) et al.
Reinforcement Learning 14
Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning
Reinforcement learning fine-tuning increasingly mixes several reward dimensions — verifiable rules, task evaluators, learned reward models — and collapses them into one scalar with fixed weights. The authors identify a failure mode where that aggregation itself causes reward hacking: static projection maps qualitatively different reward profiles onto the same scalar, pushing optimization toward whichever dimensions are easiest or densest and trapping the policy in suboptimal profiles. AMRP (Adaptive Multi-Reward Projection) reallocates aggregation weights online using relative shortfall, reward volatility, and recent progress, applying pressure to lagging or stagnant dimensions while easing off saturated ones. Across structured reasoning, citation-grounded generation, and open-ended alignment, it improves both reward-profile balance and downstream task performance over fixed and dynamic weighting baselines, and works with GRPO, GDPO, and PPO.
Group Adaptive Clipping Policy Optimization
Group-relative policy optimization for reinforcement learning with verifiable rewards (RLVR) clips the importance-sampling ratio at a fixed boundary for every rollout, which means rare correct rollouts on hard problems get suppressed at roughly the same rate as abundant correct rollouts on easy ones despite carrying much stronger exploration signal. GAPO (Group Adaptive Clipping Policy Optimization) is a drop-in change to GRPO-style methods that scales the clipping boundary with the rollout advantage, motivated by a reverse-KL trust-region argument that higher-signal rollouts deserve more update headroom; it needs no reward shaping and leaves the PPO/GSPO surrogate intact. On Qwen and Llama models it improves both Pass@1 and Pass@k over fixed clipping and advantage-shaping baselines on math and coding benchmarks where base-model pass rates are low.
GeoPAR: Large-Scale Multi-Agent Combinatorial Optimization with Geometry-Guided Parallel Autoregressive Learning
Parallel autoregressive neural solvers let multiple agents pick actions simultaneously in NP-hard routing problems, but they lose quality at scale because local geometry is modeled weakly and conflicting task selections are only repaired after actions are generated. GeoPAR adds a projection-window sparse geometry mechanism that builds lightweight local candidate neighborhoods, sparse edge-biased attention that injects those relations into node representations, and a cache-guided conflict-aware assignment step that suppresses duplicate selections during decoding rather than afterwards. On heterogeneous vehicle routing and open multi-depot pickup-and-delivery problems it improves large-scale zero-shot generalization while substantially reducing the number of rollout steps.
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning
Search and advertising retrieval systems increasingly use LLMs to expand queries, but final matching is still handed to a separate retriever. CoGR trains LLMs to generate compact keyword sets for both the query side and the item side, matched through a conventional inverted index so existing infrastructure still works; supervised fine-tuning first aligns the keyword space, then co-evolving reinforcement learning alternately optimizes each side with Group Relative Policy Optimization (GRPO) against the other's frozen index, with the item side rewarded by the counterfactual change its keywords cause in query-side F1. Against 10 sparse, dense, and generative baselines it improves F1 by 10.9% on an internal app marketplace dataset and 36.1% on the public WANDS benchmark.
CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training
Rubric-based reinforcement learning scores open-ended responses against prompt-specific checklists, but static rubrics get hacked as the policy improves, and existing dynamic rubric schemes suffer from undirected extraction, unreliable hack detection, and unbounded rubric growth. Contrastive Anchor-based Rubric Evolution (CARE) grounds each rubric update in a high-quality anchor response generated by a frontier model, contrasting the highest-scoring rollout against that anchor at every training step. An Adaptive branch reactively repairs reward misspecification while a Chase branch turns quality gaps relative to the anchor into sharper rubrics, keeping the reward discriminative in the high-reward region where over-optimization originates. Trained on WildChecklist-9K with Qwen2.5-7B base and instruct models, CARE reaches state-of-the-art results on Arena-Hard-2.0, InfoBench, and FollowBench, and is the only method whose win rate against GPT-4.1 anchor responses keeps improving across all 300 training steps, with Llama-3.1-8B-Instruct and Qwen3-8B results indicating the gains transfer across model families.
From Base Rollouts to RL Reasoning: A Budgeted Search Perspective
Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but it is unclear whether it creates capability the base model lacks or merely shifts sampling toward trajectories the base model could already reach. The authors build a Unified Decoding Framework (UDF) that expresses token sampling, beam-like search, tree search, and sequence-level resampling as policies over a shared compute budget, then test on Math500, AIME, GPQA, and IFEval whether an RL checkpoint's pass@k curve can be reproduced by a structured path of base-model operating points, using paired checkpoints from SimpleRL-Zoo. The recovery path follows a Budgeted Operating-Point Transition Rule (BOPTR), a power law relating the base and RL sampling budgets with benchmark-specific exponents, which matches RL curves on Qwen2.5-7B with a transfer error of about 3.4 percentage points and generalizes to ten models across four families and to four benchmarks it was never fitted on. The authors read this as qualified support for an internalized-search interpretation, where much of the measured RL gain under this recipe reflects improved sampling efficiency, and present the rule as a behavioral diagnostic rather than evidence of parameter-level equivalence.
Bandits in Prod: Hyperparameter Optimization at Inference Time
Many production systems can only evaluate a configuration by serving it on live traffic, which is exactly the situation for agent stacks picking a model, retrieval depth, prompting strategy, and decoding temperature without representative validation data. The authors formalize this as online hyperparameter optimization and model it as an infinitely many-armed bandit over mixed, conditional search spaces, yielding IMABO, which pairs any bandit policy over already-sampled configurations with any oracle that proposes new ones. Their restart-free anytime policy IMOSS comes with a proven expected cumulative quantile-regret bound, and paired with a Tree-structured Parzen Estimator, an incumbent-mutation oracle, or a pretrained tabular foundation model, IMABO obtains the lowest cumulative regret across settings ranging from classical machine-learning tuning to configuring LLM-based agents.
Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR
Reinforcement learning with verifiable rewards (RLVR) and benchmark scoring both hinge on an automatic verifier turning free text into a binary reward, and the known figure that one harness accepts only about 94 percent of its own gold answers is an aggregate that says nothing about which answer forms consume the error budget. The authors apply metamorphic testing to the verifiers rather than the models, generating certified meaning-preserving rewrites so every rejection is a provable false negative, then measure per-category rejection over 307,420 verdicts from four widely used verifiers. Self-validation ranges from 53.8 to 95.2 percent on identical inputs, a 41.3-point spread, with two configurations of the same library disagreeing on half of all pairs, and 93.0 percent of in-contract failures for the default LaTeX configuration tracing to whitespace and punctuation such as a trailing period or newline. Separating rejection from execution failure also shows a reference numeric cascade accepting off-by-one wrong answers as a step function of magnitude, from never below 10^4 to always at or above it, because its tolerance is relative.
Provably Safe Sim-to-Real Transfer
Policies trained in a simulator to avoid the sample cost of real-world reinforcement learning can be suboptimal once deployed, and correcting the mismatch requires real data whose collection is itself safety-constrained in domains like robotics and healthcare. Safe sim-to-real transfer is formulated within the framework of reward-free safe reinforcement learning, with a computationally efficient algorithm that exploits simulator information to provably reduce real-world interaction while keeping exploration safe and yielding a near-optimal feasible policy for any potential reward function. The real-world sample complexity bound quantifies the simulator's benefit as a function of the sim-to-real mismatch.
From Rollouts to Recipes: Self-Contained Post-Training for LLMs
Post-training normally applies a single recipe to every sample even though a model's own rollouts reveal that samples sit in different learning states. Self-Routing reads rollout correctness and confidence to send each sample to GRPO, on-policy self-distillation, regularization, or skipping, requiring no external teacher, extra annotation, or additional sampling. On mathematical reasoning with Qwen3 and Qwen3.5 backbones it consistently beats uniform GRPO, uniform on-policy self-distillation, fixed mixtures, and simpler routing baselines, and the routing distribution shifts over training as the method stops spending updates on low-signal or already stable samples.
NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games
Model-based reinforcement learning (MBRL) has worked well in single-agent settings, but extending it to two-player zero-sum imperfect-information games (IIGs) runs into opponent-induced non-stationarity and identifiability barriers that, the authors argue, make centralized model learning a mathematical necessity. NashDreamer introduces a centralized Multi-Agent Recurrent State-Space Model (MARSSM) that separates environment dynamics from the effect of each player's strategy on their own observations, and it can use any policy gradient algorithm while inheriting that algorithm's convergence guarantees toward Nash equilibria under an idealized model. On four benchmark games, it substantially improves sample efficiency over model-free baselines early in training. A theoretical analysis of the optimization landscape also identifies a vulnerability of Dreamer-style algorithms to posterior collapse in stochastic environments, which the authors leave as an open challenge.
Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers
Vision-Language Models (VLMs) carry useful priors for interactive decision-making, but using them directly as policies is expensive, does not improve with experience, and can repeat systematic errors. SAGE (Selective Agent Guidance via Entropy) queries a VLM teacher only when a lightweight reinforcement learning (RL) learner is uncertain, executes the suggested action during training, and distills the guidance into the learner's policy, optionally weighting teacher actions by environment-derived advantages because the advice is not always reliable. Across sparse-reward visual reasoning and navigation tasks, the resulting policies act without any VLM calls at evaluation time and improve over unguided RL in several environments, including settings where the learned policy exceeds its VLM teacher. Guidance helped most when the VLM steered the agent toward high-reward trajectories and least when unguided exploration already succeeded or teacher actions produced uninformative experience.
2 more specialized papers
- WiSDoM: Wireless Sparse Decision Transformer with Mixture-of-Experts for Multi-Task Mobile Network Optimization Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci
- Risk-Aware Decision-Making for Autonomous Overtaking: A World Model-Based Mixture-of-Experts Framework Yongzhi Liu, Sunan Zhang, Jinchang Xu et al.
Robotics 11
IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training
World models for embodied agents predict future frames from actions but generate physically implausible interactions, and the usual remedy injects external motion, geometry, or semantic representations that require auxiliary estimators or manual annotation. The diagnosis offered instead is a supervision-allocation mismatch in the globally averaged mean squared error denoising objective: abundant static content dominates the optimization signal, leaving the sparse dynamic regions that carry the interaction under-supervised. IMPACT uses cross-attention on manipulated-object tokens as an internal spatiotemporal prior, samples candidate regions from it, calibrates them with detached local prediction errors into an interaction map, and reweights the denoising loss accordingly, requiring no external representations and no inference-time modifications; across robot-arm and human-hand manipulation with several diffusion transformer backbones it improves interaction fidelity, physical plausibility, and visual quality over the MSE-trained baselines.
A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies
Driving policies deployed on embedded hardware are shrunk by pruning, distillation, and quantization, but the aggregate scores used to sign off on the compressed model may not reflect whether the car still drives safely around other road users. The authors train a belief-state policy with proximal policy optimization (PPO) in Gym-Duckietown, then push the extracted actor through the compression pipeline one stage at a time, evaluating in closed loop on five driving curricula after each stage. Structured pruning is the stage where driving capability first disappears; distillation partially repairs the pruned actor but is limited by its rehearsal data, and integer quantization then costs the curricula requiring the vehicle to stop and resume — whereas quantizing the unpruned actor preserves all five.
EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
Vision-language-action (VLA) models map observations and instructions to robot actions, but long-horizon tasks also require coordinating perception, planning, execution, progress verification, and recovery as the physical state changes. EmbodiedSkills treats each skill decision as an execution proposal whose prerequisites are checked before execution and whose outcome is verified afterward, connecting high-level skill selection, bounded low-level VLA execution, and post-action verification through a fixed executable-skill interface that also logs structured trajectories usable as supervision for individual components. Instantiated with Qwen3-VL and OpenPI/pi0.5, task-adapted low-level policies reach 86.20% average success across 50 RoboTwin 2.0 tasks and 97.40% across the four LIBERO suites, though the same approach achieves only 12.5% on four memory-dependent RMBench tasks.
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
How far multimodal large language models (MLLMs) extend from perceiving to acting is tested by dropping one directly into a drone's control loop with its entire action space declared solely in the prompt. DroneCATS-Agent makes the model a swappable component and the DroneCATS benchmark treats it as the independent variable across approaching a visible target, tracking a moving one, searching outside the initial view, and commanding a multi-drone fleet, with no fine-tuning or function-calling schemas and a roster scaling down to 2B parameters. Small open models often navigate into the success radius more reliably than frontier models yet lose the episode by declaring arrival prematurely or never at all, and multi-drone commanding widens the gap as small models blindly copy one coordinate across distinct views, making protocol discipline and correct termination, not navigation, the bottleneck.
Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds
Imitation-learned manipulation policies are usually stress-tested against variation in scenes, objects, or instructions rather than in how fast the task is executed, leaving open how much of the expert's speed tolerance the learner inherits. The comparison uses ParcelStow, a contact-rich task where a robot acquires, reorients, and inserts a parcel; a scripted expert and an ACT (Action Chunking with Transformers) policy trained on its demonstrations both reach 100 percent success at nominal speed. At the fastest demonstrated speed the expert still succeeds 84 percent of the time while ACT falls to 53 percent, with 35 of its 47 failures being insertion misalignments, and across all policies and speeds none of the 414 acquisitions lacking force closure ever completed the task. Matching success at nominal speed therefore says nothing about whether temporal robustness survived imitation.
Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation
Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. Facet-0 unifies multimodal representation learning and reinforcement learning (RL) post-training around a joint action-wrench proposal: a causal wrench history is aligned with vision-language semantics and kinematic state, flow matching generates each action chunk together with the wrist-wrench profile it is expected to induce, and deployment rollouts train a distributional Action-Wrench Critic that separates motions with similar task progress but different contact outcomes, with phase-aware rewards and contact-selective credit focusing improvement on decisive interactions. A lightweight bounded actor reuses the frozen representation for on-robot adaptation to part-specific dynamics while RL stays defined over executable Cartesian actions. Trained on ManuFacet-1K, a 1,000-hour force-synchronized corpus spanning three embodiments and multiple manufacturing cells, the task-adapted system reaches 82% mean success on five sub-millimeter computer-assembly tasks versus 15% for the strongest baseline, with 0.5 mm placement accuracy and 50 ms command latency.
5 more specialized papers
- Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC Baha Zarrouki, Arslan Thobani, Jasper Hoffmann et al.
- DNC-IMM: Early Lane-Change Intention Recognition via Neural Calibration Based on Driving Context Information Woong-Chan Byun, Seung-Hyun Kong
- REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs Riyaaz Shaik, Chandru Venkataraman
- Dual Process Motion Planning Jiayi Yan, Francesco Fabiano, Alessandro Abate
- Scalable Rao-Blackwellized Online Planning for High-Dimensional POMDPs Jiho Lee, Nisar Ahmed, Kyle Hollins Wray et al.
Reasoning 8
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
Rewriting benchmark problems is a common defence against data contamination in evaluating LLM mathematics, but current rewriting methods cannot guarantee that the new problem is well posed or that the new answer is right. RePro folds Lean-oriented neural automated theorem provers into the rewriting loop, regenerating both problems and answers with correctness backed by machine-checked Lean proofs. On GSM8K and MATH, the instances RePro retains hit 100% well-definedness, feasibility, and answer correctness where prior methods still emit invalid or wrongly answered items, and several models lose accuracy on the verified rewrites — a sign their scores partly reflect memorization of surface form and structure.
Dependency-Aware Chain-of-Thought Compression for Financial Reasoning
Chain-of-thought prompting helps on hard reasoning but its long intermediate traces drive up inference cost enough to obstruct deployment in financial settings. The Hierarchical Semantic Distillation Network (HSDN) compresses those chains through semantic segmentation, dependency graph construction, dual-encoder importance scoring, constrained segment selection, and local boundary rewriting, using a frozen Qwen3 4B only for feature extraction and final answer generation so the compression stage itself stays structured and interpretable. On the AFAC2025 benchmark it reaches 91.0% accuracy at 68.4% compression, ahead of strong compression baselines on overall score and reasoning coherence.
Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs
Inference-time search with language models tends to pile its budget onto a handful of structurally similar trajectories, a failure the authors name reasoning basin collapse. BASIN is a training-free selection rule that clusters reasoning states into basins and penalizes revisiting the same strategy, redirecting a fixed compute budget toward genuinely different paths; a quality-aware variant, QA-BASIN, keeps strong basins when pure diversification overshoots. Against Tree of Thoughts under matched budgets it gains up to 22 percentage points on Game of 24 and 6.7 points on MuSR, and a proposed redundancy gap statistic — how differently search concentrates on correct versus incorrect predictions — shifts from roughly zero under Tree of Thoughts to consistently positive.
DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
Self-play lets a model generate its own training questions, but without direction its solver plateaus: existing unguided signals such as difficulty, learnability, or diversity keep questions hard and varied without saying which reasoning weaknesses to attack, while guided methods import direction from human examples or document corpora outside the loop. DiagEvo derives that direction from the solver's own failure history — a diagnostician extracts recurring error causes into a hierarchical memory that groups them under skill nodes and marks each Active or Mastered by self-consistency on targeted questions, and the challenger uses those states and recurrence counts to trade off cause-targeted generation against free exploration, with double-confidence filtering keeping mid-difficulty questions only when the majority answer leads clearly. With a 4B diagnostician it beats every baseline in mean accuracy across all nine benchmarks for Qwen3-4B, Qwen3-8B, and OctoThinker-8B, reaching 72.3% mean accuracy on five mathematical reasoning benchmarks with Qwen3-8B, 4.5 points above R-Zero.
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Self-evolution methods claim to let models learn autonomously from raw material, but there has been no direct way to measure how efficiently they convert study material into problem-solving ability. StudyBench is a controlled physics benchmark that splits evaluation into an Application Set of hard textbook problems, testing absorption, and a Transfer Set of olympiad-level problems, testing generalisation. Benchmarking representative self-evolution methods on three base models shows gains on the Application Set rarely carry over to the Transfer Set, and an ablation exposes a "Guidance Gap": the strongest method recovers only a small fraction of what the same material unlocks when simply supplied as in-context guidance. Every method also plateaus well before its compute budget is exhausted, which the authors read as evidence that the bottleneck is methodological rather than a shortage of data or compute.
Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs
Chain-of-thought reasoning commits each step as text, so errors propagate and training requires traces to imitate; reasoning in a model's continuous representation space avoids these constraints but leaves open how the latent states should be computed. Latent Recurrent Thoughts (LRT) keeps a large language model (LLM) frozen as the decoder while a task-dedicated proposer supplies base latents and a tiny recurrent reasoner refines them over many steps through bounded residual corrections, decoupling depth of computation from model size. On Countdown-4, Sudoku, HumanEval, MBPP, and StrategyQA, LRT substantially outperforms prior frozen-decoder continuous-space reasoning methods under an identical decoder, prompt, data, and training budget, and it beats non-thinking-mode chain-of-thought prompting on the same backbone at a small fraction of its inference compute.
Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning
Diffusion denoisers and recursive reasoners both iterate but differ in what they carry between steps; adding a persistent hidden state to a denoiser and removing its timestep conditioning leaves a single shared update that can run to arbitrary depth. The resulting anytime solver keeps improving well past the rollout lengths and backpropagation window used in training, reaching 99.90% exact solve on Sudoku-Extreme and 98.93% on Maze-Unique. Progressive denoising turns out to be unnecessary at inference: holding corruption at maximum by replacing every non-clue variable with fresh Gaussian noise each step still converges to stable correct answers, with no parallel rollouts, candidate selection, or external verifier. Ordered annealed corruption remains essential during training, suggesting diffusion's contribution here is a denoising curriculum rather than a sampling procedure.
1 more specialized paper
- A Certificate-Producing Cascade for Equational Implication: The SAIR EQT2 Stage 2 Solver Haobo Ma, Wenlin Zhang, Manuel Israel C\'azares