Wednesday, August 12, 2026
Highlights
Training Variable Long Sequences with Data-Centric Parallel
Training on batches of widely varying sequence lengths forces a choice between static configurations that waste compute through workload imbalance and complex load-balancing systems that require invasive code changes. Data-Centric Parallel (DCP) lets the data drive the runtime instead, dynamically adjusting parallel size, gradient accumulation, and recomputation per batch based on its sequence lengths. The approach delivers up to a 2.88x speedup on 32 H200 GPUs and integrates into new models with about ten lines of code.
Training on datasets whose sequence lengths span orders of magnitude forces a bad choice: static parallel configurations waste GPUs on workload imbalance and communication overhead, while adaptive compiler-based systems demand invasive code changes. DCP (Data-Centric Parallel) instead lets each batch's sequence length drive the runtime directly, dynamically picking the sequence-parallel size, gradient accumulation, and recomputation settings per batch.
- A cheap "dual-layer" profiling pass runs just two transformer layers per (batch size, parallel size) candidate to estimate full-model speed and memory, after which
DCP-interbalances workers by giving fast batches extra gradient-accumulation steps rather than shrinking batch sizes, andDCP-intraadditionally turns off activation checkpointing for short sequences where spare memory allows. - On 32 H200 GPUs across three synthetic length distributions, DCP reaches up to 2.88× throughput over a bucket-parallel baseline (2.70× from
DCP-interalone on the 5BTransformer-1D, 1.68× on the 1.2BTransformer-2D), with the largest gains on short-sequence-dominated data. - Dropping unnecessary recomputation gives
DCP-intraa further 20–25% speedup overDCP-interfor sequences under 200k tokens, and DCP maintains a consistently low idle-time ratio and near-linear weak scaling where the baseline's throughput decays with cluster size. - The method is deliberately lightweight — integration into a new model takes at most 10 lines of code, and it composes with existing sequence-parallel schemes like
DeepSpeed UlyssesandZeRO-1. - It currently applies only to Transformer architectures and single-model systems, and the evaluation uses synthesized datasets and compares against bucket parallel rather than heavier compiler-based schedulers like
HotSPa.
The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
Internal linear probes and verbalized confidence give conflicting pictures of when a language model is about to fail, and the study measures that gap on multi-hop arithmetic chains with corrupted context. Probes detect the corrupted context with near-perfect accuracy yet are uninformative about whether the final answer is correct, while models forced into structured confidence formats collapse to two values with indistinguishable error rates, a knowing-but-not-saying pattern that holds across model families including reasoning models. As a real-time monitor, probe-triggered interventions are sharply model- and error-type-dependent: a branch-and-pick strategy is net-positive and uniquely non-breaking on Llama-3.1-8B, but reprompting and replacing prior context break correct traces at roughly the rate they rescue wrong ones.
Language models that silently operate on corrupted context give no verbal sign of it, even though the error is plainly encoded in their activations — this work shows that detecting an internal error state and predicting whether it will corrupt the final answer are two very different problems.
- On a contrastive dataset of 1,400 multi-hop arithmetic traces with five injected error types, a linear probe on the residual stream detects corrupted context with AUROC above 0.98 across five
QwenandLlamavariants (peaking at 0.997 onLlama-3.1-8B), yet the same probe's failure-prediction AUROC collapses to near chance (0.47–0.53). - Verbal channels are uninformative: elicited confidence collapses to two extreme values with indistinguishable wrong rates, no surface uncertainty signal exceeds 0.70 detection AUROC, and instruct models produce zero hedging across 258 corrupted traces each.
- Enabling chain-of-thought on
Qwen3-4Bleaves probe AUROC essentially unchanged but drops error-condition accuracy from 27.9% to 1.2%, so extended reasoning widens rather than closes the knowing-saying gap. - As a runtime monitor, probe-triggered
branch-and-pickis net-positive on every model and uniquely non-breaking onLlama-3.1-8B(4 rescued, 0 broken), whilerepromptandreplace-priorbreak correct traces at roughly the rate they rescue wrong ones. - The authors claim decodability rather than causation, report their pre-registered persistence-beats-peak hypothesis as refuted, and note that the proposed error-type-aware routing rests on small per-type samples (9–16 traces) and a classifier they do not yet have.
MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents
Evaluating embodied agents through annotated question-answering pairs is labor-intensive, while task-success metrics can hide agents that complete tasks through unsafe or nonsensical spatial behavior. Borrowing metamorphic testing from software engineering, the framework automatically generates test cases from real execution trajectories using metamorphic relations grounded in logical rules and physical laws, encoded as executable Prolog rules whose violations indicate spatial cognition failures. Across three embodied scenarios it detected 90,422 spatial cognition errors in state-of-the-art agents driven by multimodal large language models, and all tested agents scored between 0.44 and 0.52 on the proposed Spatial Cognition score against a human benchmark of 0.96.
Evaluating embodied agents by task success or hand-annotated VQA hides "false positive successes," where an agent reaches its goal despite misjudging directions, distances, or sizes along the way. MetaSpace sidesteps the test-oracle problem with metamorphic testing: it auto-generates test cases from agents' real execution trajectories and checks their spatial judgments for self-consistency against logical and physical laws, no ground-truth labels required.
- Six metamorphic relations — transitivity, symmetry, non-contradiction, the triangle inequality, size-depth consistency, and object size-ratio consistency — are encoded as executable
Prologrules, and agent answers are converted into facts so a logic engine can flag any violation as a spatial cognition error. - Running 30,300 auto-generated test cases against six MLLM-driven agents (
GPT-5,GPT-4o,Claude Sonnet 4,Qwen-VL,InternVL3.5-8B,DeepSeek-VL2-small) across household navigation, robotic-arm manipulation, and drone scenarios surfaced 90,422 spatial cognition errors. - All agents scored between 0.44 and 0.52 on the proposed Spatial Cognition score, far below the 0.96 human baseline, with directional reasoning the weakest link (below 0.38 across the board) while magnitude estimation mostly cleared 0.5.
- The errors are not marginal misses: relaxing the physical-law tolerance thresholds by 50% improved
GPT-4o's score by only about 0.05, indicating gross violations of consistency rather than borderline estimates. - Chain-of-thought prompting barely helped in mitigation experiments, whereas spatially-aware prompting with cognitive maps showed early promise; results also lean on structured answer formats and a YOLO-based object pipeline, and the size/depth relations assume an idealized pinhole-camera model.
Applications 15
15 more specialized papers
- TRIBE: Predicting Team Performance via Communication Behavior Ensembles Ali Jalal-Kamali, Nikolos Gurney, David V. Pynadath et al.
- Application of Artificial Intelligence for Fraudulent Banking Operations Recognition Bohdan Mytnyk, Oleksandr Tkachyk, Nataliya Shakhovska et al.
- Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use Yizhu Gao, Zhongzhou Chen, Min Li et al.
- Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources Sundaraparipurnan Narayanan
- Emotion in an active inference model of human driving Julian F. Schumann, Johan Engstr\"om, Ran Wei et al.
- EmoPatient: An Emotion-Directed Patient Simulator for Realistic Palliative Care Communication Training Yining Wu, Tianshu Du, Jinrui Fang et al.
- Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients Jiading Zhu, Xinyu Cindy Wang, Thomas Nguyen et al.
- PIVOT: Preference-based Intervention Vectors for Pedagogical Tutor Steering Fares Fawzi, Jiaxu Zhao, Tanya Nazaretsky et al.
- How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare Dorothee Amelung, Andrew M. Bean, Sabine C. Herpertz et al.
- EMMR: Emotion-Mediated Multimodal Reasoning for Personality Assessment in Asynchronous Video Interviews Dongsheng Hu, Tianyi Zhang, Chuang Liu et al.
- Representation Matters in Longitudinal Affective Computing Igor Matias, Maximilian Haas, Eric J. Daza et al.
- KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service Assistants Saurabh Sakalkar, Abhishek Singh, Ramesh Raskar
- From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations Hongfei Yan, Jiangkai Xiong, Yiqing Li et al.
- WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management Yi Zhang, Hongyang Wang, Zheng Hao Leong et al.
- NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation Yuchen Zhou, Niels Bobet, Maribel Acosta
Agents 7
From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents
Research agents that run multi-round machine-learning experiments in industrial recommendation systems retain trajectories to guide later decisions, but those trajectories can contain unsupported artifacts, invalid or confounded rounds, and findings obscured by later edits. The proposed framework converts trajectories into evidence through a context-isolated generate-verify-repair loop that checks artifacts before release, followed by post-execution validity and attribution checks that qualify each claim as an actionable repair, a diagnostic guard, or a withheld finding, preserved as auditable records with explicit provenance; an LLM-assisted controller then applies, defers, or rejects records against new targets. Audits expose non-monotonic trajectory evolution, with final rounds frequently underperforming an earlier best, and candidates from the full workflow achieved positive online lifts over deployed baselines.
Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents
Game benchmarks for language agents mostly score final outcomes, revealing little about how players learn from repeated interaction. The study formulates experience-sensitive game learning as a framework for analyzing behavioral change across repeated play, introducing a suite of interactive games with reusable strategic structure, cross-game greedy-to-global metrics, and behavioral diagnostics computed from action traces, applied to both human players and recent self-evolving language agents. Humans show interpretable, stable shifts from locally greedy heuristics toward global strategies, while self-evolving agents show noisy and transient gains, suggesting current self-evolution methods struggle to convert gameplay experience into durable changes in decision-making.
Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation
Whether persona-prompted large language models can reproduce individual users' social media reactions matters both for stress-testing recommender systems and for gauging the threat of synthetic agent swarms to public opinion. Twelve LLM configurations are benchmarked on binary like/dislike prediction across 296 survey-based agent profiles and 26 ground-truth-mapped posts at three levels of profile completeness, against leave-post-out machine-learning baselines. GPT-5.5 Pro reaches 96.68% accuracy with full profiles but drops to chance level with demographics alone, showing that demographic inference carries negligible signal, while supervised classifiers collapse to 15.4% on unseen posts where LLMs sustain zero-shot generalization. The least heterogeneous configuration also homogenizes 34% of simulated population reactions, underscoring both the validation and the risk.
DocAtlas: Long-Document Understanding as Mutable-State Interaction
Answering questions over long documents with many pages, tables, and figures usually means either retrieving from a static index before generation or prompting a frozen proprietary model with tools. DocAtlas instead wraps the model in a mutable document harness: an external environment exposing search, reading, note-taking, and review tools that maintains a hierarchical tree and note store updated as the agent records evidence, all under a fixed context budget. With GPT-5.4 the system reaches 71.4% on MMLongBench-Doc, exceeding the 65.8% human-expert reference, and a Qwen3.5-4B vision-language model trained with end-to-end reinforcement learning in the same environment reaches 63.7% versus a 54.4% direct-input baseline.
Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
Outcome rewards for search-augmented agents cannot tell grounded retrieval apart from redundant searching, while richer process supervision or LLM judges are expensive to run during training. Search-G1 derives intrinsic rewards from the policy's own representations using two intervention-calibrated readouts: a prompt-state readout predicting whether closed-book knowledge suffices, and an answer-commit readout measuring how sensitive the answer is to deleting retrieved evidence, with both refit periodically as reinforcement learning shifts the policy. Across search-based question-answering benchmarks and two model scales, the approach improves the grounding versus search-cost trade-off, yielding shorter trajectories at competitive accuracy without process annotations or judge inference during training.
MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents
Evaluating embodied agents through annotated question-answering pairs is labor-intensive, while task-success metrics can hide agents that complete tasks through unsafe or nonsensical spatial behavior. Borrowing metamorphic testing from software engineering, the framework automatically generates test cases from real execution trajectories using metamorphic relations grounded in logical rules and physical laws, encoded as executable Prolog rules whose violations indicate spatial cognition failures. Across three embodied scenarios it detected 90,422 spatial cognition errors in state-of-the-art agents driven by multimodal large language models, and all tested agents scored between 0.44 and 0.52 on the proposed Spatial Cognition score against a human benchmark of 0.96.
1 more specialized paper
Other 5
5 more specialized papers
- Coordinated incentives in AI-generated misinformation governance Qin Li, Gui Zhang, Minyu Feng et al.
- Towards an Argumentative Foundation for Evaluative AI Xiang Yin, Tim Miller, Nico Potyka et al.
- Harnessing Abundance: A Generativity Perspective on Human-GenAI Collaboration Yoram M Kalman, Yun Wan
- Innovating with Generative AI: A Human Bottleneck Framework Julian De Freitas, Ayelet Israeli, Gideon Nave et al.
- An evolutionary model of animats with VLM-based subjective evaluation Shota Miyazaki, Takaya Arita, Reiji Suzuki
Safety & Alignment 5
Divergent Response Modes in Frontier Language Models Under Steering Pressure
Frontier language models are trained with different data, objectives, and safety pipelines, but whether this yields measurably different behavior under explicit steering has been underexplored. Six frontier models from six developers are evaluated on 300 paired base and steered prompts spanning values conflicts, reasoning elicitation, and reasoning suppression, with all six models acting as blind peer judges and 24,480 judgments scored by leave-one-out consensus. Models differ not just in how much steering shifts behavior but in the kind of response they produce: GPT-5 deflects requests to disclose its reasoning while leaving answers intact 99% of the time, versus 0% for every other model, and Claude Opus 4.7 and GPT-5 resist suppression instructions in distinct ways. Using Llama as the open-weight model, a linear probe decodes the largest behavioral split from the residual stream at 0.87 held-out accuracy, and injecting that direction during generation drives the behavior from 0% to 86%.
The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
Internal linear probes and verbalized confidence give conflicting pictures of when a language model is about to fail, and the study measures that gap on multi-hop arithmetic chains with corrupted context. Probes detect the corrupted context with near-perfect accuracy yet are uninformative about whether the final answer is correct, while models forced into structured confidence formats collapse to two values with indistinguishable error rates, a knowing-but-not-saying pattern that holds across model families including reasoning models. As a real-time monitor, probe-triggered interventions are sharply model- and error-type-dependent: a branch-and-pick strategy is net-positive and uniquely non-breaking on Llama-3.1-8B, but reprompting and replacing prior context break correct traces at roughly the rate they rescue wrong ones.
Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
Fusing multiple modalities reshapes the safety landscape of large models in ways that frameworks built on uni-modal assumptions do not capture. The survey proposes a multimodal-grounded taxonomy of threats covering adversarial attacks, data poisoning, jailbreaks, and hallucinations, arguing that modality integration, misalignment, and fusion create qualitatively new attack surfaces that shift the underlying threat models. It then organizes recent defense strategies around updated safety assumptions and closes with open challenges for building principled, scalable safety mechanisms for multimodal systems.
2 more specialized papers
- Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains Hiroki Naito
- JaleesBench: Are AI Assistants Good Spiritual Company? M. Waleed Kadous (iaser.ai, Faith Family Technology Network), Benjamin Olsen (Faith Family Technology Network)
Theory 4
Quantifying the noise sensitivity of the Wasserstein metric for images
Wasserstein (optimal transport) distances are increasingly used to compare images, but their behavior under pixel-wise additive noise has lacked theoretical grounding. Treating images as discrete measures on the pixel grid, the authors derive finite-sample expectation bounds under a Gaussian noise model, proving that the error in the signed 2-Wasserstein discrepancy scales with the square root of the noise standard deviation, versus linear scaling for the Euclidean metric. Experiments support the bounds, uncover a counterintuitive effect where adding noise can decrease the Wasserstein distance, and a cryo-electron microscopy case study shows the metric capturing data-manifold geometry at noise levels where the Euclidean metric fails.
Optimized Sequential Testing for Binary Ensemble Classifiers
Ensemble classifiers such as random forests gain accuracy by combining many base models, but evaluating every base model at prediction time is costly. Borrowing from sequential testing, the method evaluates base models one at a time and halts once a clear majority emerges, formalizing three notions of optimal early stopping that each reduce to an efficiently solvable linear program constrained by an allowable disagreement rate with the full ensemble. On datasets from the UC Irvine Machine Learning Repository and the Grinsztajn et al. tabular benchmarks, the optimal strategies deliver speed-ups of 4x or more while holding disagreement to 0.1%.
2 more specialized papers
Large Language Models 3
Training Variable Long Sequences with Data-Centric Parallel
Training on batches of widely varying sequence lengths forces a choice between static configurations that waste compute through workload imbalance and complex load-balancing systems that require invasive code changes. Data-Centric Parallel (DCP) lets the data drive the runtime instead, dynamically adjusting parallel size, gradient accumulation, and recomputation per batch based on its sequence lengths. The approach delivers up to a 2.88x speedup on 32 H200 GPUs and integrates into new models with about ten lines of code.
2 more specialized papers
- Cross-Model Humor Preference Modeling with Cards Against Humanity Victor Winter, Farhan Lakhany
- How to Ask the AI: A User Perspective Survey for Large Language Model Prompting Yiqun Zhang, Yunfan Zhang, Mingjie Zhao et al.