Wednesday, August 12, 2026

40 papers cs.AI · cs.LG · cs.CL 2026-08-13 →

Jul Aug Sep

Highlights

Training Variable Long Sequences with Data-Centric Parallel

Highlight Large Language Models Geng Zhang, Xuanlei Zhao, Kai Wang, Yang You Training on batches of widely varying sequence lengths forces a choice between static configurations that waste compute through workload imbalance and complex load-balancing systems that require invasive code changes. Data-Centric Parallel (DCP) lets the data drive the runtime instead, dynamically adjusting parallel size, gradient accumulation, and recomputation per batch based on its sequence lengths. The approach delivers up to a 2.88x speedup on 32 H200 GPUs and integrates into new models with about ten lines of code.

Training on datasets whose sequence lengths span orders of magnitude forces a bad choice: static parallel configurations waste GPUs on workload imbalance and communication overhead, while adaptive compiler-based systems demand invasive code changes. DCP (Data-Centric Parallel) instead lets each batch's sequence length drive the runtime directly, dynamically picking the sequence-parallel size, gradient accumulation, and recomputation settings per batch.

  • A cheap "dual-layer" profiling pass runs just two transformer layers per (batch size, parallel size) candidate to estimate full-model speed and memory, after which DCP-inter balances workers by giving fast batches extra gradient-accumulation steps rather than shrinking batch sizes, and DCP-intra additionally turns off activation checkpointing for short sequences where spare memory allows.
  • On 32 H200 GPUs across three synthetic length distributions, DCP reaches up to 2.88× throughput over a bucket-parallel baseline (2.70× from DCP-inter alone on the 5B Transformer-1D, 1.68× on the 1.2B Transformer-2D), with the largest gains on short-sequence-dominated data.
  • Dropping unnecessary recomputation gives DCP-intra a further 20–25% speedup over DCP-inter for sequences under 200k tokens, and DCP maintains a consistently low idle-time ratio and near-linear weak scaling where the baseline's throughput decays with cluster size.
  • The method is deliberately lightweight — integration into a new model takes at most 10 lines of code, and it composes with existing sequence-parallel schemes like DeepSpeed Ulysses and ZeRO-1.
  • It currently applies only to Transformer architectures and single-model systems, and the evaluation uses synthesized datasets and compares against bucket parallel rather than heavier compiler-based schedulers like HotSPa.

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

Highlight Safety & Alignment Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk Internal linear probes and verbalized confidence give conflicting pictures of when a language model is about to fail, and the study measures that gap on multi-hop arithmetic chains with corrupted context. Probes detect the corrupted context with near-perfect accuracy yet are uninformative about whether the final answer is correct, while models forced into structured confidence formats collapse to two values with indistinguishable error rates, a knowing-but-not-saying pattern that holds across model families including reasoning models. As a real-time monitor, probe-triggered interventions are sharply model- and error-type-dependent: a branch-and-pick strategy is net-positive and uniquely non-breaking on Llama-3.1-8B, but reprompting and replacing prior context break correct traces at roughly the rate they rescue wrong ones.

Language models that silently operate on corrupted context give no verbal sign of it, even though the error is plainly encoded in their activations — this work shows that detecting an internal error state and predicting whether it will corrupt the final answer are two very different problems.

  • On a contrastive dataset of 1,400 multi-hop arithmetic traces with five injected error types, a linear probe on the residual stream detects corrupted context with AUROC above 0.98 across five Qwen and Llama variants (peaking at 0.997 on Llama-3.1-8B), yet the same probe's failure-prediction AUROC collapses to near chance (0.47–0.53).
  • Verbal channels are uninformative: elicited confidence collapses to two extreme values with indistinguishable wrong rates, no surface uncertainty signal exceeds 0.70 detection AUROC, and instruct models produce zero hedging across 258 corrupted traces each.
  • Enabling chain-of-thought on Qwen3-4B leaves probe AUROC essentially unchanged but drops error-condition accuracy from 27.9% to 1.2%, so extended reasoning widens rather than closes the knowing-saying gap.
  • As a runtime monitor, probe-triggered branch-and-pick is net-positive on every model and uniquely non-breaking on Llama-3.1-8B (4 rescued, 0 broken), while reprompt and replace-prior break correct traces at roughly the rate they rescue wrong ones.
  • The authors claim decodability rather than causation, report their pre-registered persistence-beats-peak hypothesis as refuted, and note that the proposed error-type-aware routing rests on small per-type samples (9–16 traces) and a classifier they do not yet have.

MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

Highlight Agents Gengyang Xu, Dongwei Xiao, Yiteng Peng, Shuai Wang Evaluating embodied agents through annotated question-answering pairs is labor-intensive, while task-success metrics can hide agents that complete tasks through unsafe or nonsensical spatial behavior. Borrowing metamorphic testing from software engineering, the framework automatically generates test cases from real execution trajectories using metamorphic relations grounded in logical rules and physical laws, encoded as executable Prolog rules whose violations indicate spatial cognition failures. Across three embodied scenarios it detected 90,422 spatial cognition errors in state-of-the-art agents driven by multimodal large language models, and all tested agents scored between 0.44 and 0.52 on the proposed Spatial Cognition score against a human benchmark of 0.96.

Evaluating embodied agents by task success or hand-annotated VQA hides "false positive successes," where an agent reaches its goal despite misjudging directions, distances, or sizes along the way. MetaSpace sidesteps the test-oracle problem with metamorphic testing: it auto-generates test cases from agents' real execution trajectories and checks their spatial judgments for self-consistency against logical and physical laws, no ground-truth labels required.

  • Six metamorphic relations — transitivity, symmetry, non-contradiction, the triangle inequality, size-depth consistency, and object size-ratio consistency — are encoded as executable Prolog rules, and agent answers are converted into facts so a logic engine can flag any violation as a spatial cognition error.
  • Running 30,300 auto-generated test cases against six MLLM-driven agents (GPT-5, GPT-4o, Claude Sonnet 4, Qwen-VL, InternVL3.5-8B, DeepSeek-VL2-small) across household navigation, robotic-arm manipulation, and drone scenarios surfaced 90,422 spatial cognition errors.
  • All agents scored between 0.44 and 0.52 on the proposed Spatial Cognition score, far below the 0.96 human baseline, with directional reasoning the weakest link (below 0.38 across the board) while magnitude estimation mostly cleared 0.5.
  • The errors are not marginal misses: relaxing the physical-law tolerance thresholds by 50% improved GPT-4o's score by only about 0.05, indicating gross violations of consistency rather than borderline estimates.
  • Chain-of-thought prompting barely helped in mitigation experiments, whereas spatially-aware prompting with cognitive maps showed early promise; results also lean on structured answer formats and a YOLO-based object pipeline, and the size/depth relations assume an idealized pinhole-camera model.

Applications 15

15 more specialized papers

Agents 7

From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents

Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He et al. cross-listed Research agents that run multi-round machine-learning experiments in industrial recommendation systems retain trajectories to guide later decisions, but those trajectories can contain unsupported artifacts, invalid or confounded rounds, and findings obscured by later edits. The proposed framework converts trajectories into evidence through a context-isolated generate-verify-repair loop that checks artifacts before release, followed by post-execution validity and attribution checks that qualify each claim as an actionable repair, a diagnostic guard, or a withheld finding, preserved as auditable records with explicit provenance; an LLM-assisted controller then applies, defers, or rejects records against new targets. Audits expose non-monotonic trajectory evolution, with final rounds frequently underperforming an earlier best, and candidates from the full workflow achieved positive online lifts over deployed baselines.

Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents

Yingying Guo, Zhuoxuan Ju, Ruibo Ming, Ruicheng Feng, Jinjin Gu cross-listed Game benchmarks for language agents mostly score final outcomes, revealing little about how players learn from repeated interaction. The study formulates experience-sensitive game learning as a framework for analyzing behavioral change across repeated play, introducing a suite of interactive games with reusable strategic structure, cross-game greedy-to-global metrics, and behavioral diagnostics computed from action traces, applied to both human players and recent self-evolving language agents. Humans show interpretable, stable shifts from locally greedy heuristics toward global strategies, while self-evolving agents show noisy and transient gains, suggesting current self-evolution methods struggle to convert gameplay experience into durable changes in decision-making.

Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation

Ljubisa Bojic, Ljiljana Matic, Joerg Matthes, Milan Cabarkapa, Bojana Dinic, Jue Wang cross-listed Whether persona-prompted large language models can reproduce individual users' social media reactions matters both for stress-testing recommender systems and for gauging the threat of synthetic agent swarms to public opinion. Twelve LLM configurations are benchmarked on binary like/dislike prediction across 296 survey-based agent profiles and 26 ground-truth-mapped posts at three levels of profile completeness, against leave-post-out machine-learning baselines. GPT-5.5 Pro reaches 96.68% accuracy with full profiles but drops to chance level with demographics alone, showing that demographic inference carries negligible signal, while supervised classifiers collapse to 15.4% on unseen posts where LLMs sustain zero-shot generalization. The least heterogeneous configuration also homogenizes 34% of simulated population reactions, underscoring both the validation and the risk.

DocAtlas: Long-Document Understanding as Mutable-State Interaction

Hongchen Wei, Yuanzhe Wang, Bei Liu, Yifan Yang, Qi Dai, Kai Qiu et al. cross-listed Answering questions over long documents with many pages, tables, and figures usually means either retrieving from a static index before generation or prompting a frozen proprietary model with tools. DocAtlas instead wraps the model in a mutable document harness: an external environment exposing search, reading, note-taking, and review tools that maintains a hierarchical tree and note store updated as the agent records evidence, all under a fixed context budget. With GPT-5.4 the system reaches 71.4% on MMLongBench-Doc, exceeding the 65.8% human-expert reference, and a Qwen3.5-4B vision-language model trained with end-to-end reinforcement learning in the same environment reaches 63.7% versus a 54.4% direct-input baseline.

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

Cheng Ruoxi, Ma Haoxuan, Zhang Hongyi, Zhang Junming, Duan Ranjie, Xia Qiaolin et al. cross-listed Outcome rewards for search-augmented agents cannot tell grounded retrieval apart from redundant searching, while richer process supervision or LLM judges are expensive to run during training. Search-G1 derives intrinsic rewards from the policy's own representations using two intervention-calibrated readouts: a prompt-state readout predicting whether closed-book knowledge suffices, and an answer-commit readout measuring how sensitive the answer is to deleting retrieved evidence, with both refit periodically as reinforcement learning shifts the policy. Across search-based question-answering benchmarks and two model scales, the approach improves the grounding versus search-cost trade-off, yielding shorter trajectories at competitive accuracy without process annotations or judge inference during training.

MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

Gengyang Xu, Dongwei Xiao, Yiteng Peng, Shuai Wang Evaluating embodied agents through annotated question-answering pairs is labor-intensive, while task-success metrics can hide agents that complete tasks through unsafe or nonsensical spatial behavior. Borrowing metamorphic testing from software engineering, the framework automatically generates test cases from real execution trajectories using metamorphic relations grounded in logical rules and physical laws, encoded as executable Prolog rules whose violations indicate spatial cognition failures. Across three embodied scenarios it detected 90,422 spatial cognition errors in state-of-the-art agents driven by multimodal large language models, and all tested agents scored between 0.44 and 0.52 on the proposed Spatial Cognition score against a human benchmark of 0.96.
1 more specialized paper

Other 5

5 more specialized papers

Safety & Alignment 5

Divergent Response Modes in Frontier Language Models Under Steering Pressure

Ali Jalal-Kamali cross-listed Frontier language models are trained with different data, objectives, and safety pipelines, but whether this yields measurably different behavior under explicit steering has been underexplored. Six frontier models from six developers are evaluated on 300 paired base and steered prompts spanning values conflicts, reasoning elicitation, and reasoning suppression, with all six models acting as blind peer judges and 24,480 judgments scored by leave-one-out consensus. Models differ not just in how much steering shifts behavior but in the kind of response they produce: GPT-5 deflects requests to disclose its reasoning while leaving answers intact 99% of the time, versus 0% for every other model, and Claude Opus 4.7 and GPT-5 resist suppression instructions in distinct ways. Using Llama as the open-weight model, a linear probe decodes the largest behavioral split from the residual stream at 0.87 held-out accuracy, and injecting that direction during generation drives the behavior from 0% to 86%.

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk Internal linear probes and verbalized confidence give conflicting pictures of when a language model is about to fail, and the study measures that gap on multi-hop arithmetic chains with corrupted context. Probes detect the corrupted context with near-perfect accuracy yet are uninformative about whether the final answer is correct, while models forced into structured confidence formats collapse to two values with indistinguishable error rates, a knowing-but-not-saying pattern that holds across model families including reasoning models. As a real-time monitor, probe-triggered interventions are sharply model- and error-type-dependent: a branch-and-pick strategy is net-positive and uniquely non-breaking on Llama-3.1-8B, but reprompting and replacing prior context break correct traces at roughly the rate they rescue wrong ones.

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu, Yusen Zhang et al. cross-listed Fusing multiple modalities reshapes the safety landscape of large models in ways that frameworks built on uni-modal assumptions do not capture. The survey proposes a multimodal-grounded taxonomy of threats covering adversarial attacks, data poisoning, jailbreaks, and hallucinations, arguing that modality integration, misalignment, and fusion create qualitatively new attack surfaces that shift the underlying threat models. It then organizes recent defense strategies around updated safety assumptions and closes with open challenges for building principled, scalable safety mechanisms for multimodal systems.
2 more specialized papers

Theory 4

Quantifying the noise sensitivity of the Wasserstein metric for images

Erik Lager, Gilles Mordant, Amit Moscovich cross-listed Wasserstein (optimal transport) distances are increasingly used to compare images, but their behavior under pixel-wise additive noise has lacked theoretical grounding. Treating images as discrete measures on the pixel grid, the authors derive finite-sample expectation bounds under a Gaussian noise model, proving that the error in the signed 2-Wasserstein discrepancy scales with the square root of the noise standard deviation, versus linear scaling for the Euclidean metric. Experiments support the bounds, uncover a counterintuitive effect where adding noise can decrease the Wasserstein distance, and a cryo-electron microscopy case study shows the metric capturing data-manifold geometry at noise levels where the Euclidean metric fails.

Optimized Sequential Testing for Binary Ensemble Classifiers

Joseph Kalman, Amit Moscovich cross-listed Ensemble classifiers such as random forests gain accuracy by combining many base models, but evaluating every base model at prediction time is costly. Borrowing from sequential testing, the method evaluates base models one at a time and halts once a clear majority emerges, formalizing three notions of optimal early stopping that each reduce to an efficiently solvable linear program constrained by an allowable disagreement rate with the full ensemble. On datasets from the UC Irvine Machine Learning Repository and the Grinsztajn et al. tabular benchmarks, the optimal strategies deliver speed-ups of 4x or more while holding disagreement to 0.1%.
2 more specialized papers

Large Language Models 3

Training Variable Long Sequences with Data-Centric Parallel

Geng Zhang, Xuanlei Zhao, Kai Wang, Yang You Training on batches of widely varying sequence lengths forces a choice between static configurations that waste compute through workload imbalance and complex load-balancing systems that require invasive code changes. Data-Centric Parallel (DCP) lets the data drive the runtime instead, dynamically adjusting parallel size, gradient accumulation, and recomputation per batch based on its sequence lengths. The approach delivers up to a 2.88x speedup on 32 H200 GPUs and integrates into new models with about ten lines of code.
2 more specialized papers

Multimodal 1

Unified Hallucination Fuzzing for Multimodal Large Language Models

Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si et al. cross-listed Static hallucination benchmarks for multimodal large language models cover narrow failure taxonomies and saturate quickly, masking how brittle models are on evolving inputs. The framework pairs UniHall, a fine-grained benchmark spanning Object, Instruction, and Knowledge dimensions, with SAMF (Self-Adaptive Multimodal Fuzzing), which uses evolutionary mutation strategies and an ensemble of multimodal oracles to probe hallucination boundaries dynamically. State-of-the-art models degrade significantly under fuzzing compared with conventional evaluation, exposing a dissociation between reasoning ability and factual grounding, and the experiments surface a helpfulness-hallucination trade-off in which reinforcement-learning alignment worsens sycophancy on instruction-following tasks.