Do LLMs Understand Limit Order Book Dynamics?
arXiv:2608.23706v1 Announce Type: new Abstract: A large language model (LLM) trained on synthetic limit order book (LOB) data achieves near perfect s
arXiv:2608.23706v1 Announce Type: new Abstract: A large language model (LLM) trained on synthetic limit order book (LOB) data achieves near perfect s
arXiv:2608.23740v1 Announce Type: new Abstract: Concurrent multi-agent coding promises division of labor across modules, robustness through redundanc
arXiv:2608.23807v1 Announce Type: new Abstract: Masked diffusion language models (dLLMs) can in principle generate text faster than autoregressive (A
arXiv:2608.23811v1 Announce Type: new Abstract: Automated fact-checking is essential for ensuring the reliability of public health information, yet t
arXiv:2608.23817v1 Announce Type: new Abstract: SHAP and LIME are now standard tools for interpreting black-box predictions, yet their outputs can va
arXiv:2608.23834v1 Announce Type: new Abstract: The key-value (KV) cache is a primary capacity and bandwidth bottleneck in long-context LLM serving.
arXiv:2608.23837v1 Announce Type: new Abstract: Large language models (LLMs) are known to exhibit social sycophancy, often validating or agreeing wit
arXiv:2608.23848v1 Announce Type: new Abstract: Budget-constrained agentic search arises when an LLM agent must refine candidates under a small evalu
arXiv:2608.23855v1 Announce Type: new Abstract: We propose ICI-Time, a novel framework that reframes time series forecasting as a visual inpainting t
arXiv:2608.23870v1 Announce Type: new Abstract: When it comes to safety policies for generative AI, one size does not fit all. Each organization and
arXiv:2608.23873v1 Announce Type: new Abstract: Everything a language model sees is tokens. The serving stack knows what each span is -- user input,
arXiv:2608.23875v2 Announce Type: new Abstract: Artificial Intelligence (AI) algorithms frequently learn creative and unexpected solutions, surprisin
arXiv:2608.23893v1 Announce Type: new Abstract: Learning systems deployed over long periods must adapt not only to statistical changes in incoming da
arXiv:2608.23898v1 Announce Type: new Abstract: We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification
arXiv:2608.23906v1 Announce Type: new Abstract: Artificial Intelligence (AI) is increasingly integrated into complex sociotechnical systems, includin
arXiv:2608.23908v1 Announce Type: new Abstract: Tax-loss harvesting demonstrates consistent benefits to long-term portfolio growth; yet implementing
arXiv:2608.23911v1 Announce Type: new Abstract: Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling t
arXiv:2608.23918v1 Announce Type: new Abstract: Large Language Models excel at code generation, yet competitive programming exposes a persistent fail
arXiv:2608.23922v1 Announce Type: new Abstract: Data mixing is a central design problem in large language model pretraining: given a fixed token budg
arXiv:2608.23932v1 Announce Type: new Abstract: This study introduces the evolutionarily recurrent decision model (ERDM), a computational reinforceme
arXiv:2608.23941v1 Announce Type: new Abstract: Pre-execution oversight is core to trusted monitoring in AI control: a fallible LLM monitor vets plan
arXiv:2608.23956v1 Announce Type: new Abstract: Test-time reasoning methods such as iterative refinement, decomposition, and repeated sampling are of
arXiv:2608.23962v1 Announce Type: new Abstract: When an LLM serving deployment runs out of KVcache room, there are two well-established ways out. Ten
arXiv:2608.23970v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made significant progress in understanding and interpre
arXiv:2608.23978v1 Announce Type: new Abstract: Visual grounding is typically evaluated as a one-shot mapping from an informative referring expressio
arXiv:2608.23979v1 Announce Type: new Abstract: In a deliberative poll, once submissions outnumber what anyone will read, some mechanism chooses whic
arXiv:2608.23982v1 Announce Type: new Abstract: Scientific reasoning requires language models to retrieve specialized knowledge and incorporate it re
arXiv:2608.24001v2 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for future prediction, motivating the use of multi
arXiv:2608.24005v1 Announce Type: new Abstract: Knowledge Tracing (KT) aims to assess students' dynamic knowledge states from their learning historie
arXiv:2608.24015v1 Announce Type: new Abstract: The Planner-Operator-Reflector (POR) framework is widely used in GUI agents to maintain objective ali
arXiv:2608.24024v1 Announce Type: new Abstract: Confidence-based voting aggregates parallel LLM rollouts by weighting each with internal signals such
arXiv:2608.24041v1 Announce Type: new Abstract: Although Speech Large Language Models (SpeechLLMs) excel at speech understanding and generation, thei
arXiv:2608.24046v1 Announce Type: new Abstract: When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem
arXiv:2608.24069v1 Announce Type: new Abstract: LLM-based multi-agent trading systems, in which specialized agents collaborate through structured com
arXiv:2608.24070v1 Announce Type: new Abstract: Prohibitive computational and environmental costs impede the scalable deployment of Large Language Mo
arXiv:2608.24076v2 Announce Type: new Abstract: Evaluation of agentic information retrieval remains limited to scripted interactions with uniform use
arXiv:2608.24086v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as code agents for scientific and engineering anal
arXiv:2608.24099v1 Announce Type: new Abstract: GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-up
arXiv:2608.24103v1 Announce Type: new Abstract: Commercial design platforms increasingly edit documents through large language model (LLM) agents, bu
arXiv:2608.24112v1 Announce Type: new Abstract: Modern text-to-image (T2I) models often have similar total scores but different strengths, making pra
arXiv:2608.24114v1 Announce Type: new Abstract: Training multi-turn LLM agents with reinforcement learning typically relies on trajectory-level rewar
arXiv:2608.24135v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has emerged as a pivotal technique for enhancin
arXiv:2608.24160v1 Announce Type: new Abstract: Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and t
arXiv:2608.24174v1 Announce Type: new Abstract: Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outc
arXiv:2608.24188v1 Announce Type: new Abstract: Coding agents re-send large file reads and tool outputs to a frontier LLM every turn, and this contex
arXiv:2608.24192v1 Announce Type: new Abstract: Aligning large language models to human preferences is crucial for real-world deployment but frequent
arXiv:2608.24214v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) requires language models to decide when to continue sear
arXiv:2608.24218v1 Announce Type: new Abstract: Enterprise entity alignment must handle semi-structured records, implicit attributes, and unit or gra
arXiv:2608.24228v1 Announce Type: new Abstract: Many LLM applications are most useful when they provide several candidate outputs for comparison, val
arXiv:2608.24232v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content,
arXiv:2608.24237v1 Announce Type: new Abstract: Longitudinal radiology report generation (LRRG) requires identifying both current findings and their
arXiv:2608.24252v1 Announce Type: new Abstract: LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful implemen
arXiv:2608.24258v1 Announce Type: new Abstract: AI systems are increasingly evaluated for legally accountable settings, where correct outputs must al
arXiv:2608.24263v1 Announce Type: new Abstract: Change data synthesis provides a cost-effective solution for expanding training data and improving th
arXiv:2608.24273v2 Announce Type: new Abstract: Continual knowledge graph embedding updates entity and relation representations as a graph grows. Exi
arXiv:2608.24274v1 Announce Type: new Abstract: A sustainable diet represents a multi-dimensional synergy among four essential pillars: nutrition ade
arXiv:2608.24275v1 Announce Type: new Abstract: Safeguarding language model agents requires assessing complete execution trajectories under context-d
arXiv:2608.24291v1 Announce Type: new Abstract: Paper-to-code reproduction asks scientific AI agents to turn research papers into executable reposito
arXiv:2608.24302v1 Announce Type: new Abstract: Long-video understanding depends critically on how a limited model context is constructed from a much
arXiv:2608.24310v1 Announce Type: new Abstract: Search-augmented reasoning remains difficult for small language models. On-policy distillation (OPD)
arXiv:2608.24314v1 Announce Type: new Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture b
arXiv:2608.24319v1 Announce Type: new Abstract: An intelligent system does not merely reason: it governs its own reasoning - how much to compute, whe
arXiv:2608.24325v1 Announce Type: new Abstract: Reliable underwater perception requires complementary sensing under variable visibility. Optical came
arXiv:2608.24338v1 Announce Type: new Abstract: Inference-time decoding methods improve LLM reasoning by exploring multiple candidate trajectories, y
arXiv:2608.24358v1 Announce Type: new Abstract: Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. A
arXiv:2608.24361v1 Announce Type: new Abstract: Multi-agent LLM systems are increasingly deployed in real-world applications, where failures can be c
arXiv:2608.24368v1 Announce Type: new Abstract: Reliable multi-turn tool use requires an agent to preserve an evolving task state and ensure that eac
arXiv:2608.24369v1 Announce Type: new Abstract: While large language models (LLMs) possess vast zero-shot procedural knowledge, their tendency to pro
arXiv:2608.24411v1 Announce Type: new Abstract: The efficiency of Large Language Model (LLM) serving is fundamentally limited by the sequential natur
arXiv:2608.24419v1 Announce Type: new Abstract: LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, b
arXiv:2608.24427v1 Announce Type: new Abstract: Non-parametric (partial) identification of counterfactual queries typically relies on a fully specifi
arXiv:2608.24441v1 Announce Type: new Abstract: Electric vehicle (EV) charging loads exhibit strong behavioral heterogeneity and temporal variability
arXiv:2608.24462v1 Announce Type: new Abstract: In this paper, we propose \textbf{Mahalanobis-Based Multi-Head Attention} (MHA-CSP), a novel attentio
arXiv:2608.24467v1 Announce Type: new Abstract: Although recent Multimodal Large Language Models (MLLMs) have advanced general product understanding,
arXiv:2608.24470v1 Announce Type: new Abstract: Heterogeneous agile Earth observation satellite (AEOS) scheduling requires task selection, satellite
arXiv:2608.24471v1 Announce Type: new Abstract: Maritime moving-target observation scheduling with agile Earth observation satellites is a dynamic, s
arXiv:2608.24509v1 Announce Type: new Abstract: LLM agents increasingly solve tasks by invoking multiple tools, where parallel execution is essential
arXiv:2608.24534v1 Announce Type: new Abstract: Clinical LLMs can generate recommendations that are factually plausible yet physiologically unsafe. W
arXiv:2608.24545v1 Announce Type: new Abstract: Human collective intelligence depends on transmission processes: who shares what with whom, how, and
arXiv:2608.24569v1 Announce Type: new Abstract: Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflo
arXiv:2608.24570v1 Announce Type: new Abstract: Clinical diagnosis is an active evidence-seeking process in which clinicians acquire evidence, update
arXiv:2608.24571v1 Announce Type: new Abstract: Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creati
arXiv:2608.24574v1 Announce Type: new Abstract: Video multimodal large language models support language guided video segmentation, but they often sho
arXiv:2608.24585v1 Announce Type: new Abstract: Automated high-density storage systems (warehouses, robotic parking, plant logistics, etc.) require f
arXiv:2608.24632v1 Announce Type: new Abstract: Accurate assessment of student competencies is essential for enabling educators to identify individua
arXiv:2608.24658v1 Announce Type: new Abstract: Scaling test-time reasoning has substantially improved the problem-solving ability of large language
arXiv:2608.24662v2 Announce Type: new Abstract: Evaluations of generative language models frequently interpret observable behavioral traits, such as
arXiv:2608.24691v1 Announce Type: new Abstract: Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidenc
arXiv:2608.24713v1 Announce Type: new Abstract: Lifted inference algorithms enable scalable probabilistic inference even for large object domains by
arXiv:2608.24735v1 Announce Type: new Abstract: Self-improving LLM agents refine answers, not the process that produces those answers. Systems that a
arXiv:2608.24758v1 Announce Type: new Abstract: Discovering stable neuron behavior across entire domains remains a challenge in mechanistic interpret
arXiv:2608.24764v1 Announce Type: new Abstract: Large language model agents are moving beyond conventional retrieval-augmented generation toward dire
arXiv:2608.24777v1 Announce Type: new Abstract: LLM-based agents can interact with external environments through tool invocation, but this capability
arXiv:2608.24790v1 Announce Type: new Abstract: Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the v
arXiv:2608.24794v1 Announce Type: new Abstract: Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neithe
arXiv:2608.24804v1 Announce Type: new Abstract: We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping m
arXiv:2608.24810v1 Announce Type: new Abstract: Recent work has applied Mamba style state space models (SSMs) to video anomaly detection, yet existin
arXiv:2608.24824v1 Announce Type: new Abstract: Large language models are increasingly used for knowledge graph question answering (KGQA), but can fa
arXiv:2608.24825v1 Announce Type: new Abstract: The rapid expansion of large-scale assessments and the growing adoption of automatic item generation
arXiv:2608.24846v1 Announce Type: new Abstract: Real-world data for knowledge graph question answering is often distributed across different organiza
arXiv:2608.24870v1 Announce Type: new Abstract: Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly
arXiv:2608.24876v1 Announce Type: new Abstract: Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure
arXiv:2608.23258v1 Announce Type: cross Abstract: We propose HetSkills, a novel framework designed to progressively learn heterogeneous skills within
arXiv:2608.23572v1 Announce Type: cross Abstract: We propose a novel framework for the analysis of multimodal data -- encompassing visual, auditory,
arXiv:2608.23593v1 Announce Type: cross Abstract: Text-to-image systems use learned aesthetic scorers to filter training data and guide generation, b
arXiv:2608.23611v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer new opportunities for automated code refactoring. However, gener
arXiv:2608.23616v2 Announce Type: cross Abstract: An AI agent's rebuild is only as good as the process that produced it. Prior work found that once a
arXiv:2608.23619v1 Announce Type: cross Abstract: Legacy software repositories embed decades of domain knowledge in undocumented code, making underst
arXiv:2608.23623v1 Announce Type: cross Abstract: Tool-using agents must decide when to stop. Existing systems already gate terminal success, certify
arXiv:2608.23629v1 Announce Type: cross Abstract: Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Pla
arXiv:2608.23635v1 Announce Type: cross Abstract: Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them
arXiv:2608.23651v1 Announce Type: cross Abstract: Agent harnesses record a failed tool call and its error message in the transcript and ask the model
arXiv:2608.23653v1 Announce Type: cross Abstract: AI agents are increasingly used for simulation-driven engineering. Physical system modeling present
arXiv:2608.23658v1 Announce Type: cross Abstract: An LLM serving engine sizes its key-value (KV) cache once, at startup, permanently setting aside a
arXiv:2608.23660v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural
arXiv:2608.23663v2 Announce Type: cross Abstract: Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device
arXiv:2608.23705v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly capable of generating text that challenges human perf
arXiv:2608.23752v1 Announce Type: cross Abstract: The growing size of Convolutional Neural Networks has led to increasingly large and costly models.
arXiv:2608.23758v1 Announce Type: cross Abstract: Recent large audio language models (LALMs) have achieved impressive progress in audio understanding
arXiv:2608.23763v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model
arXiv:2608.23766v1 Announce Type: cross Abstract: Between AI-assisted item generation and expert review sits a computational evaluator whose decision
arXiv:2608.23776v1 Announce Type: cross Abstract: Understanding human skill is important for AI systems that collaborate with, coach, or assist peopl
arXiv:2608.23780v1 Announce Type: cross Abstract: LLMs are being used increasingly to measure aspects of student discourse (e.g. talk moves, collabor
arXiv:2608.23791v1 Announce Type: cross Abstract: Psychological research on emotion dynamics has established that human affect is a continuous, evolv
arXiv:2608.23799v1 Announce Type: cross Abstract: Recent progress in image restoration has converged on all-in-one architectures that jointly handle
arXiv:2608.23803v1 Announce Type: cross Abstract: Lung cancer tissue diagnostics is complex, as therapy decisions in precision oncology rely on the i
arXiv:2608.23809v1 Announce Type: cross Abstract: Multilingual language models can solve the same mathematical problem in different languages, but it
arXiv:2608.23814v1 Announce Type: cross Abstract: Large Language Models (LLMs) demonstrate strong capabilities in automated essay scoring (AES), but
arXiv:2608.23824v1 Announce Type: cross Abstract: Unmanned aerial vehicle (UAV)-mounted 5G New Radio base stations (gNBs) can augment terrestrial net
arXiv:2608.23836v1 Announce Type: cross Abstract: Accurate interpretation of volumetric CT requires efficient navigation of 3D image volumes and atte
arXiv:2608.23838v1 Announce Type: cross Abstract: Healthcare documentation in the neonatal intensive care unit (NICU) presents significant challenges
arXiv:2608.23839v1 Announce Type: cross Abstract: Embodied Agents System (EAS) are increasingly deployed in open-world physical domains, where reliab
arXiv:2608.23840v1 Announce Type: cross Abstract: Training large-scale AI models often outgrows a single data center, demanding sharded, multi-cluste
arXiv:2608.23842v1 Announce Type: cross Abstract: DevOps programming (e.g., using CLI/API scripts or IaC frameworks) is key to cloud infrastructure m
arXiv:2608.23847v1 Announce Type: cross Abstract: This paper proposes the Coronavirus Optimization Algorithm (COA), a SARS-CoV-2-inspired success-his
arXiv:2608.23858v1 Announce Type: cross Abstract: The Agent Payments Protocol (AP2), introduced by Google, enables large language model (LLM)-driven
arXiv:2608.23860v1 Announce Type: cross Abstract: Revelation Control is the problem of choosing priced interventions that reveal hidden state only in
arXiv:2608.23885v1 Announce Type: cross Abstract: Real-time optimization (RTO) relies on process models to locate economically optimal operating cond
arXiv:2608.23892v1 Announce Type: cross Abstract: This article presents the abridged core of \emph{A Mathematical Theory of Interpretation} (MTI), wh
arXiv:2608.23895v1 Announce Type: cross Abstract: Kohn--Sham density functional theory (DFT) underpins electronic-structure simulations, but repeated
arXiv:2608.23897v1 Announce Type: cross Abstract: When a code generating language model fabricates a Python package name, an adversary who has pre-re
arXiv:2608.23928v1 Announce Type: cross Abstract: Surgical spatio-temporal grounding (STG) requires locating, at each video time specified by a proce
arXiv:2608.23934v1 Announce Type: cross Abstract: Nitrogen-vacancy (NV) centers in diamond can serve as highly sensitive solid-state quantum sensors
arXiv:2608.23943v1 Announce Type: cross Abstract: High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and a
arXiv:2608.23952v1 Announce Type: cross Abstract: Federated video anomaly detection trains model collaboratively without sharing raw surveillance foo
arXiv:2608.23953v1 Announce Type: cross Abstract: An agent harness is what turns a language model into an autonomous agent: the surrounding code that
arXiv:2608.23959v1 Announce Type: cross Abstract: Safety alignment in large language models (LLMs) remains brittle against a growing spectrum of atta
arXiv:2608.23961v1 Announce Type: cross Abstract: Background: Large Language Models (LLMs) have demonstrated strong performance across a variety of c
arXiv:2608.23965v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves the factuality of large language models by grounding
arXiv:2608.23986v1 Announce Type: cross Abstract: Large language model providers are compute constrained, and their universal response to congestion
arXiv:2608.23992v1 Announce Type: cross Abstract: Large language model (LLM) agents invoke external tools to retrieve and reason over information bey
arXiv:2608.24011v1 Announce Type: cross Abstract: Chinese ancient document understanding demands complex visual, linguistic, and historical reasoning
arXiv:2608.24017v1 Announce Type: cross Abstract: The emerging W3C WebMCP proposal enables LLM agents to invoke tools exposed by web pages. In multi-
arXiv:2608.24020v1 Announce Type: cross Abstract: Generating executable parametric CAD code from dimension-annotated orthographic drawings is a chall
arXiv:2608.24022v1 Announce Type: cross Abstract: LLM agents integrated with external resources gain complex task capabilities, yet the unified natur
arXiv:2608.24033v1 Announce Type: cross Abstract: Time series classification underpins applications in healthcare, sensing, and industrial monitoring
arXiv:2608.24039v1 Announce Type: cross Abstract: Manufacturing process planning transforms heterogeneous design information into coherent manufactur
arXiv:2608.24042v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models pretrained on large-scale robot datasets provide a strong
arXiv:2608.24048v1 Announce Type: cross Abstract: While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific qu
arXiv:2608.24063v1 Announce Type: cross Abstract: While Vision Large Language Models (VLLMs) have achieved remarkable success in multimodal reasoning
arXiv:2608.24065v1 Announce Type: cross Abstract: While recent advances in data synthesis aim to curate high-quality datasets, most generation pipeli
arXiv:2608.24073v1 Announce Type: cross Abstract: Low-earth-orbit (LEO) satellites enable high-resolution, large-scale Earth observation for applicat
arXiv:2608.24080v1 Announce Type: cross Abstract: In psychological counseling, effective support is not always delivered through long, information-ri
arXiv:2608.24082v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown strong capabilities in table reasoning, but their effective
arXiv:2608.24087v1 Announce Type: cross Abstract: Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or aft
arXiv:2608.24107v1 Announce Type: cross Abstract: Material replacement is a common interior-design operation: changing the material of a selected sur
arXiv:2608.24113v1 Announce Type: cross Abstract: Time-series anomalies can appear not only as pointwise deviations but also as changes in recurring
arXiv:2608.24115v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial
arXiv:2608.24119v1 Announce Type: cross Abstract: Visual demonstrations provide a natural interface for specifying image transformations that are dif
arXiv:2608.24130v1 Announce Type: cross Abstract: Multi-camera 3D perception systems for warehouse scenes are trained largely on synthetic data and e
arXiv:2608.24132v1 Announce Type: cross Abstract: Product catalogs in fast-moving service businesses are shifting from static, independently priced S
arXiv:2608.24133v1 Announce Type: cross Abstract: People search for urban outdoor places not only by category or function, but also by what activitie
arXiv:2608.24154v1 Announce Type: cross Abstract: Real-world deployment of traffic surveillance systems is bottlenecked by geographic domain shift, i
arXiv:2608.24156v1 Announce Type: cross Abstract: Industrial actor--critic methods usually represent continuous actions as anonymous numerical coordi
arXiv:2608.24163v1 Announce Type: cross Abstract: Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive T
arXiv:2608.24176v1 Announce Type: cross Abstract: Item tokenizer encodes semantic embeddings into token IDs to replace the randomly assigned item IDs
arXiv:2608.24191v1 Announce Type: cross Abstract: Urdu, the world's tenth most spoken language with 246 million speakers, remains almost entirely abs
arXiv:2608.24300v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) enables language models to learn multi-turn i
arXiv:2608.24304v1 Announce Type: cross Abstract: Recent controllable text generation (CTG) for sentiment control has largely focused on decoder-base
arXiv:2608.24340v1 Announce Type: cross Abstract: The prediction of student engagement from the online tutoring videos is difficult because engagemen
arXiv:2608.24342v1 Announce Type: cross Abstract: Synthetic image generation is a promising strategy to address data scarcity and the underrepresenta
arXiv:2608.24350v1 Announce Type: cross Abstract: To reduce the hallucination risk caused by outcome-driven rewards in large language models trained
arXiv:2608.24354v1 Announce Type: cross Abstract: MLLMs are increasingly deployed in user-facing applications, yet they inherit backdoor risks from t
arXiv:2608.24384v1 Announce Type: cross Abstract: Resistance training can be a high risk activity, and safe form is essential to avoiding injury. Lab
arXiv:2608.24386v1 Announce Type: cross Abstract: Tensor-valued prediction is fundamental to geometric deep learning, yet uncertainty quantification
arXiv:2608.24400v1 Announce Type: cross Abstract: We study multilevel fair resource allocation with tree-structured hierarchical relations among agen
arXiv:2608.24436v1 Announce Type: cross Abstract: Wearable device data enables continuous health monitoring, but suffers from structured missingness:
arXiv:2608.24482v1 Announce Type: cross Abstract: Mechanistic Localization bridges mechanistic interpretability and post-training optimization by iso
arXiv:2608.24492v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) methods are widely used for hallucination detection in large langua
arXiv:2608.24500v1 Announce Type: cross Abstract: The STAR (Student-Teacher Achievement Ratio) experiment (1985, Tennessee, USA) is a landmark hierar
arXiv:2608.24524v1 Announce Type: cross Abstract: Feature attribution is a central tool of model interpretability, yet the software through which it
arXiv:2608.24551v1 Announce Type: cross Abstract: Machine learning models are widely used in financial fraud and credit-risk detection, yet their adv
arXiv:2608.24555v1 Announce Type: cross Abstract: Prehospital stroke assessment aims to accurately identify stroke symptoms and make rapid decisions
arXiv:2608.24559v1 Announce Type: cross Abstract: Despite the critical role of grey literature in scholarly communication, artefacts such as Calls fo
arXiv:2608.24568v1 Announce Type: cross Abstract: Deep neural networks generalize well despite their highly nonconvex, overparameterized loss landsca
arXiv:2608.24582v1 Announce Type: cross Abstract: Credit risk models increasingly need to combine predictive accuracy with transparent explanations a
arXiv:2608.24597v1 Announce Type: cross Abstract: Electroencephalography (EEG) is a widely used window into human brain function, but most EEG models
arXiv:2608.24644v1 Announce Type: cross Abstract: This paper introduces an environment for constructing literate programs in concert with language-aw
arXiv:2608.24650v2 Announce Type: cross Abstract: System-level simulation is an essential tool for exploring the rapidly expanding design space of LL
arXiv:2608.24664v1 Announce Type: cross Abstract: We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 an
arXiv:2608.24696v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become
arXiv:2608.24712v1 Announce Type: cross Abstract: Optimization of hyperparameters is a critical factor to obtain optimal model performance. While exi
arXiv:2608.24715v1 Announce Type: cross Abstract: A vast amount of optical satellite data is being transmitted to Earth-based servers every day, and
arXiv:2608.24721v1 Announce Type: cross Abstract: Hyperparameter selection remains a key challenge in Bayesian optimization (BO) and Bayesian active
arXiv:2608.24727v1 Announce Type: cross Abstract: EEG foundation models pretrained via self-supervised learning promise transferable representations,
arXiv:2608.24748v1 Announce Type: cross Abstract: How can humans make sense of the rapid takeoff of artificial intelligence (AI)? We studied the sens
arXiv:2608.24753v1 Announce Type: cross Abstract: Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end corr
arXiv:2608.24762v1 Announce Type: cross Abstract: Many disentanglement methods represent generative factors using Euclidean product coordinates, alth
arXiv:2608.24768v1 Announce Type: cross Abstract: The Bayesian Ideal Observer (IO) establishes the theoretical upper bound on task performance for bi
arXiv:2608.24771v1 Announce Type: cross Abstract: Brain stroke, known for its high mortality and incidence rates, poses significant health risks and
arXiv:2608.24807v1 Announce Type: cross Abstract: Model cards are structured documents that summarize key information about machine learning models t
arXiv:2608.24842v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosu
arXiv:2608.24845v1 Announce Type: cross Abstract: We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B
arXiv:2201.13427v2 Announce Type: replace Abstract: This article discusses a particular case of the data clustering problem, where it is necessary to
arXiv:2304.10041v3 Announce Type: replace Abstract: This work investigates formal policy synthesis for continuous-state stochastic dynamic systems su
arXiv:2312.17535v2 Announce Type: replace Abstract: In the past two years, the outstanding performance of ChatGPT in multilingual and multitasking ha
arXiv:2406.05375v4 Announce Type: replace Abstract: Root cause analysis (RCA) is crucial for enhancing the reliability and performance of complex sys
arXiv:2506.11578v5 Announce Type: replace Abstract: Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple
arXiv:2507.11482v5 Announce Type: replace Abstract: Artificial learning systems are graduating from passive learners to increasingly autonomous agent
arXiv:2511.02605v3 Announce Type: replace Abstract: Shielding is widely used to enforce safety in reinforcement learning (RL), ensuring that an agent
arXiv:2511.08873v4 Announce Type: replace Abstract: Large language models (LLMs) are shifting from answer providers to intelligent tutors in educatio
arXiv:2512.13979v2 Announce Type: replace Abstract: Large reasoning models achieve strong performance on diverse tasks by producing extended chains o
arXiv:2601.10485v5 Announce Type: replace Abstract: Domain-specific knowledge graphs (DKGs) are critical yet often suffer from limited coverage compa
arXiv:2602.02304v3 Announce Type: replace Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as s
arXiv:2602.09159v2 Announce Type: replace Abstract: Recent multi-agent frameworks have shown promise for oncology decision support, yet most assume c
arXiv:2604.01532v3 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Pr
arXiv:2604.01841v4 Announce Type: replace Abstract: Clinical prediction from structured electronic health records (EHRs) is challenging due to high d
arXiv:2604.15994v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) excel at recognizing individual visual elements and reas
arXiv:2605.05535v2 Announce Type: replace Abstract: The evaluation of housing potential requires consideration of a location from multiple perspectiv
arXiv:2605.10059v3 Announce Type: replace Abstract: Agent-based modeling (ABM) has long been used in economics to study human behavior, and large lan
arXiv:2605.19743v3 Announce Type: replace Abstract: Engineering-agent systems are proliferating, but differences in tasks, tools, and success criteri
arXiv:2606.08405v3 Announce Type: replace Abstract: While data-intensive deep reinforcement learning can optimize complex control policies, scientifi
arXiv:2607.12634v2 Announce Type: replace Abstract: This paper proposes a theoretical and empirical framework for understanding intelligence as a pro
arXiv:2608.08253v2 Announce Type: replace Abstract: We present SuperLocalMemory 4.0, a governed, local-first memory operating system for AI agents, u
arXiv:2608.09696v4 Announce Type: replace Abstract: A primary goal of science is to learn mechanistic or causal world models from data. These model
arXiv:2608.14426v2 Announce Type: replace Abstract: AI is increasingly being used to help with AI R&D. Under certain conditions this feedback loo
arXiv:2608.14673v2 Announce Type: replace Abstract: We present an independent human assessment of the proof developed in Chapter 6 of OpenAI's Ten Ad
arXiv:2608.16645v3 Announce Type: replace Abstract: Can a language model recover the true research idea of a published paper when given only that pap
arXiv:2608.20009v2 Announce Type: replace Abstract: Understanding object dynamics requires not only predicting future trajectories but also examining
arXiv:2608.20054v3 Announce Type: replace Abstract: Multi-module neural systems often expose every module to the full input. We test whether a slot-s
arXiv:2608.22018v2 Announce Type: replace Abstract: Hate speech research has moved from coarse-grained classification towards structured parsing, whe
arXiv:2608.22055v2 Announce Type: replace Abstract: Suppose one embodied agent knows what must be built, while its teammate alone knows which transfo
arXiv:2608.22232v2 Announce Type: replace Abstract: Real-world situation appearances can deviate from their underlying physical states, challenging t
arXiv:2608.22577v2 Announce Type: replace Abstract: Long-horizon GUI agents can retain complete action histories as compact text, but only a few hist
arXiv:2608.22979v2 Announce Type: replace Abstract: Deploying high-dimensional multimodal features in industrial recommender systems incurs substanti
arXiv:2608.23035v2 Announce Type: replace Abstract: As on-device LLM agents evolve into personal copilots, the mobile operating system has become a k
arXiv:2608.23283v2 Announce Type: replace Abstract: General-purpose language models can reason and synthesize knowledge, but complex work also requir
arXiv:2608.23397v2 Announce Type: replace Abstract: Interactive clinical agents operate under partial observability, so reliable care depends on reac
arXiv:2306.14300v2 Announce Type: replace-cross Abstract: Autism spectrum disorder (ASD) is a developmental condition that presents significant chall
arXiv:2402.01767v4 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) significantly improves document-based question answeri
arXiv:2407.00500v2 Announce Type: replace-cross Abstract: Recent point-based intrinsic decomposition and inverse rendering methods have advanced the
arXiv:2412.02520v4 Announce Type: replace-cross Abstract: Connected automated vehicles (CAVs) equipped with adaptive cruise control (ACC) create new
arXiv:2503.17894v3 Announce Type: replace-cross Abstract: We propose generative learner for estimating heterogeneous treatment effects and characteri
arXiv:2504.18346v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have been transformative across many domains. However, halluci
arXiv:2505.23197v5 Announce Type: replace-cross Abstract: Path planning for autonomous robots faces a fundamental trade-off between path length and o
arXiv:2506.12202v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often call external tools to solve tasks. One effective strate
arXiv:2506.13182v3 Announce Type: replace-cross Abstract: [...] Since then, various APR approaches, especially those leveraging the power of large la
arXiv:2506.20073v3 Announce Type: replace-cross Abstract: Spatio-temporal data mining plays a pivotal role in informed decision making across diverse
arXiv:2507.21790v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are becoming widely used to support various workflows across d
arXiv:2508.09473v2 Announce Type: replace-cross Abstract: Ensuring robust safety alignment while preserving utility is critical for the reliable depl
arXiv:2509.01479v4 Announce Type: replace-cross Abstract: Explainable systems expose information about why certain observed effects are happening to
arXiv:2509.03754v2 Announce Type: replace-cross Abstract: Responding to rising global food security needs, precision agriculture and deep learning-ba
arXiv:2509.15959v2 Announce Type: replace-cross Abstract: Autonomous navigation in maritime domains is accelerating alongside advances in artificial
arXiv:2509.18778v2 Announce Type: replace-cross Abstract: Visual imitation learning frameworks allow robots to learn manipulation skills from expert
arXiv:2510.14249v2 Announce Type: replace-cross Abstract: Understanding and modeling the relationship between language and sound are essential for ap
arXiv:2510.23634v5 Announce Type: replace-cross Abstract: Motivated by applications for set containment problems, we consider the following fundament
arXiv:2511.01870v3 Announce Type: replace-cross Abstract: Studying the cellular architecture of the human cerebral cortex is essential for understand
arXiv:2512.12703v2 Announce Type: replace-cross Abstract: Extracting human motion from large-scale web videos offers a scalable solution to the data
arXiv:2512.16715v3 Announce Type: replace-cross Abstract: In recent years, Predictive Process Mining (PPM) techniques based on artificial neural netw
arXiv:2601.00282v3 Announce Type: replace-cross Abstract: Quantization is widely used to accelerate inference and streamline the deployment of large
arXiv:2601.07737v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in mainst
arXiv:2601.10034v3 Announce Type: replace-cross Abstract: Decision making often exhibits context dependence that is difficult to accommodate within a
arXiv:2601.16520v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual recogn
arXiv:2601.17027v2 Announce Type: replace-cross Abstract: While synthetic data has proven effective for improving scientific reasoning in the text do
arXiv:2601.19435v2 Announce Type: replace-cross Abstract: Sustainable monetization of large language models (LLMs) remains a critical open challenge.
arXiv:2602.03702v2 Announce Type: replace-cross Abstract: Large language models are increasingly trained in continual or open-ended settings, where t
arXiv:2602.10863v2 Announce Type: replace-cross Abstract: Long-horizon reinforcement learning for information seeking agents remains difficult becaus
arXiv:2602.11684v2 Announce Type: replace-cross Abstract: As Large Language Models increasingly power role-playing applications, simulating patients
arXiv:2602.13940v3 Announce Type: replace-cross Abstract: Tokenization is a hardcoded compression step which remains in the training pipeline of Larg
arXiv:2602.18532v3 Announce Type: replace-cross Abstract: Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged
arXiv:2603.00188v3 Announce Type: replace-cross Abstract: Training-free KV cache compression is essential for deploying vision-language GUI agents un
arXiv:2603.02041v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are predominantly trained on English-centric data, resulting i
arXiv:2603.10068v2 Announce Type: replace-cross Abstract: Most adversarial evaluations of large language model (LLM) safety assess single prompts and
arXiv:2603.16497v3 Announce Type: replace-cross Abstract: Time series foundation models (TSFMs) require diverse, real-world datasets to adapt across
arXiv:2603.16654v3 Announce Type: replace-cross Abstract: Evaluating the reasoning abilities of large language models (LLMs) solely from final answer
arXiv:2603.17170v2 Announce Type: replace-cross Abstract: AI agents increasingly execute users' natural-language (NL) tasks by calling Web services,
arXiv:2603.25507v2 Announce Type: replace-cross Abstract: Network Traffic Classification (NTC) increasingly relies on data-driven models, yet its pra
arXiv:2604.14211v2 Announce Type: replace-cross Abstract: This thesis is an exposition of Ollivier-Ricci Curvature of metric spaces as introduced by
arXiv:2604.16926v3 Announce Type: replace-cross Abstract: Electroencephalography (EEG) foundation models have shown strong potential for learning gen
arXiv:2605.00901v2 Announce Type: replace-cross Abstract: The use of CT imaging is important for screening, diagnosis, therapy planning, and prognosi
arXiv:2605.02867v3 Announce Type: replace-cross Abstract: Despite significant advances in Reinforcement Learning (RL), model performance remains high
arXiv:2605.06647v3 Announce Type: replace-cross Abstract: Retrieval-augmented agents are increasingly the interface to large knowledge bases, yet mos
arXiv:2605.09477v2 Announce Type: replace-cross Abstract: Methods based on diffusion models (DMs) for solving inverse problems (IPs) have recently ac
arXiv:2605.10426v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end aut
arXiv:2605.11048v2 Announce Type: replace-cross Abstract: Existing imitation learning methods enable robots to interact autonomously with the physica
arXiv:2605.26958v2 Announce Type: replace-cross Abstract: Reinforcement learning in open-ended long-form generation is challenging because reliable r
arXiv:2605.28791v2 Announce Type: replace-cross Abstract: On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged in
arXiv:2606.00930v2 Announce Type: replace-cross Abstract: Mechanistic interpretability routinely reads a probe and labels its top-activating units as
arXiv:2606.04120v2 Announce Type: replace-cross Abstract: Conversational agents that serve as lifelong companions must maintain persistent memory acr
arXiv:2606.13705v2 Announce Type: replace-cross Abstract: The Gemma 4 instruction-tuned models share a reproducible failure: on long factual enumerat
arXiv:2606.18699v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet t
arXiv:2606.22027v4 Announce Type: replace-cross Abstract: Reinforcement learning for robot manipulation is often bottlenecked by reward design, espec
arXiv:2606.24192v2 Announce Type: replace-cross Abstract: Unlearning has emerged as a key technique to mitigate harmful content generation in diffusi
arXiv:2606.29717v3 Announce Type: replace-cross Abstract: Predicting a material's properties from its structure is a central, fast-advancing problem
arXiv:2607.04028v2 Announce Type: replace-cross Abstract: We propose a unified algebraic framework for classification performance evaluation covering
arXiv:2607.08960v2 Announce Type: replace-cross Abstract: Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode compl
arXiv:2607.13393v3 Announce Type: replace-cross Abstract: A field can reformulate its computations freely exactly where its demand is stated independ
arXiv:2607.13431v2 Announce Type: replace-cross Abstract: Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternativ
arXiv:2607.19261v4 Announce Type: replace-cross Abstract: Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, exa
arXiv:2607.21580v2 Announce Type: replace-cross Abstract: Controllable video generation remains challenging due to the difficulty of specifying preci
arXiv:2607.24555v2 Announce Type: replace-cross Abstract: Serving large language models at long context is bottlenecked by the key-value (KV) cache,
arXiv:2607.29221v2 Announce Type: replace-cross Abstract: We address the challenge of securely and efficiently outsourcing AI computations from a tru
arXiv:2608.01400v2 Announce Type: replace-cross Abstract: Tabular foundation models, driven by in-context learning, have rapidly grown in quality and
arXiv:2608.06165v3 Announce Type: replace-cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the applicati
arXiv:2608.07557v2 Announce Type: replace-cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reacti
arXiv:2608.08882v4 Announce Type: replace-cross Abstract: AI tools that help people judge online claims are usually evaluated while the tool is prese
arXiv:2608.14773v2 Announce Type: replace-cross Abstract: The efficient-KAN literature---covering Chebyshev, wavelet, and radial-basis-function varia
arXiv:2608.15299v2 Announce Type: replace-cross Abstract: Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of
arXiv:2608.18076v2 Announce Type: replace-cross Abstract: Large-scale image generation has benefited from advances in data scale, quality, rebalancin
arXiv:2608.18445v3 Announce Type: replace-cross Abstract: We present the first mechanised formalisation of Romanov's Triplet Logic (TLS) in the Rocq
arXiv:2608.19760v2 Announce Type: replace-cross Abstract: Audited against policy-conditional ground truth from executed replay in a single-agent tool
arXiv:2608.20539v2 Announce Type: replace-cross Abstract: Digital twin simulations show promise, but current empirical evidence suggests that the app
arXiv:2608.20804v2 Announce Type: replace-cross Abstract: Temporal Knowledge Graph (TKG) extrapolation seeks to infer future facts from time-varying
arXiv:2608.20818v3 Announce Type: replace-cross Abstract: The matrix-aware optimizer Muon improves large model training by balancing updates across s
arXiv:2608.21114v2 Announce Type: replace-cross Abstract: Visual world-model agents such as DreamerV3 act through a recurrent latent state rather tha
arXiv:2608.21170v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent wor
arXiv:2608.21829v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation treats the document store as a frozen input, and the offline
arXiv:2608.22067v3 Announce Type: replace-cross Abstract: World-Action Models (WAMs) build robot control on video-generation backbones, which jointly
arXiv:2608.22354v2 Announce Type: replace-cross Abstract: Delta-Rule recurrent models maintain a fixed-size state, enabling $O(1)$ inference memory b
arXiv:2608.22462v2 Announce Type: replace-cross Abstract: Neural networks can acquire new capabilities while damaging existing ones, but what determi
arXiv:2608.22876v2 Announce Type: replace-cross Abstract: Hybrid sequence models must satisfy prefix invariance: representations at position t must n
arXiv:2608.23104v2 Announce Type: replace-cross Abstract: Molecular science represents an important frontier for LLM-based agents. Unlike general age
arXiv:2608.23286v2 Announce Type: replace-cross Abstract: Federated learning on non-IID data seeks flat minima to generalize across clients, and exis
arXiv:2608.23329v2 Announce Type: replace-cross Abstract: Open-world video understanding often requires a model to locate sparse visual evidence and
arXiv:2608.23391v2 Announce Type: replace-cross Abstract: Structured data exists in many forms (tables, knowledge graphs, charts, and time series), a
arXiv:2608.23474v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) achieve strong performance on video and image-sequence benchm
arXiv:2608.23566v2 Announce Type: replace-cross Abstract: Group-based reinforcement learning methods such as GRPO for large language models avoid tra
Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with di
Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer when solving a problem. As models
The development of 0.1^{circ} global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-reso
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. Libri
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents be
Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generat
Document retrieval increasingly supports high-stakes information access in finance, healthcare, and law. Modern retrieval pipelines vary both in modal
Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems. While exist
Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reaso
Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, fol
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this d
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference effi
Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution remains unstable. We first identi
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame predicti
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We
Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserv
Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for
Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, langu
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a s
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) expe
Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single unsupported object, chart value, o
Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable eva
GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to action misuse, yet existing benchmarks lack
Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segme
Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergen
Modern software -- from plugin systems to self-evolving agent harnesses -- increasingly requires dynamic composition, yet its formal foundations remai
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated progr
Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each ro
World models aim to simulate how complex environments evolve under actions and events, yet existing video-based world models primarily learn dynamics
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be
Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model
Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existing multi-agent pipelines distribu
Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream state is repeatedly transformed into
We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families purs
Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may injec
Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel exploration at the natural granul
LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runt
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task
Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedur
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment inform
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substanti
As large language models (LLMs) continue to advance in coding capabilities, their potential in cybersecurity has drawn increasing research attention,
Video games provide a scalable source of training data for video world models, offering diverse environments, complex interactions, and abundant in-th
Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, post-training on large multi-task da
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent con
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from Common
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space
Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementation in the Python ecosystem, i.e