WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data· 可穿戴健康推理基准
Recent advances in wearable sensing enable continuous monitoring of physiological and beha…
Recent advances in wearable sensing enable continuous monitoring of physiological and beha…
Does a simplified legal clause still say what the original said? The checks in current use…
Computational phylogenetics has become an essential tool in historical linguistics, yet it…
Large language models (LLMs) show strong reasoning ability, but their explanations can rem…
Assessing the impacts of social policy changes is a widely acknowledged challenge for poli…
Many recurring text functions are easy to describe but difficult to implement with rules, …
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appen…
Reasoning traces from chain-of-thought models appear to offer a legible window into how a …
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge dur…
For scientific progress, we need benchmarks that test the limits of state-of-the-art model…
Harnessing naturally occurring feedback from user interactions offers a promising learning…
Sign language processing systems have traditionally operated at the sentence level, ignori…
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibi…
Low-resource authorship style transfer (LAST) aims to rewrite text into the style of an ar…
Training data attribution (TDA) aims to identify training examples that shape model behavi…
LLM-based evaluators of natural language generation (NLG) quality are widely deployed as s…
Dynamic agent harnesses let language models change the software that shapes their own exec…
Natural language is emerging as a primary feedback channel for improving language agents, …
AI tutors are most useful when they adapt to each student's strengths, weaknesses, and pre…
Extracting structured fields from hundreds of millions of documents annually remains costl…
While WhisperX accelerates speech transcription via intra-audio batching, it isolates audi…
BioMedRAG introduced retrieval-augmented generation with a learned chunk scorer for biomed…
Large language models (LLMs) offer promising clinical decision support but remain vulnerab…
Research planning is the decisive capability of AI scientists. Yet a research plan admits …
Many important forms of human learning begin with a vague goal, such as "become a better p…
Can a listener recover what a speaker means from the form of an utterance alone? We answer…
Forced alignment evaluation typically requires manually annotated timestamps, limiting lar…
Two forms of test-time scaling for Large Language Models (LLMs) have emerged as effective …
Recent advances in large language models (LLMs) have demonstrated strong capabilities in n…
Factual question answering (QA) typically assumes a single canonical answer, obscuring whe…
Recent advances in inference-time scaling have significantly improved the reasoning perfor…
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy …
Transduced language models (TLMs) compose a pretrained \emph{source} language model with a…
Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning…
Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of la…
Test-time scaling uses extra test-time compute to improve performance, such as letting lan…
Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in re…
When a speaker refers to a scene that the listener cannot directly see, the listener must …
Multimodal instruction-following models require training data that is accurate, diverse, v…
Linguistic meaning is grounded in conceptual content, from which reference to particular e…
Web agents that act from rendered pixels avoid the fragility and heavy token cost of readi…
Large language models (LLMs) are increasingly deployed as AI analysts to process financial…
We present Crase, a bounded and inspectable alternative to deep research agents for schola…
Distinguishing machine-generated text (MGT) from human-written text (HWT) becomes increasi…
Text-to-CAD aims to generate executable CAD programs from natural-language descriptions. H…
Modern software systems accumulate technical debt over decades of development, which makes…
Recent advances in continuous diffusion and flow-based language models (LMs) have achieved…
Historical people may appear under different languages, scripts, and transcription traditi…
Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising cha…
Vision-language models (VLMs) achieve strong performance on video and image-sequence bench…
Users increasingly turn to large language models for emotional support, yet little is know…
That a prompt's effect is not a property of the prompt is established: prompts optimised f…
Large language models often rely on Chain-of-Thought (CoT) reasoning to solve complex task…
Question answering (QA) over long, connected documents remains challenging because relevan…
While recent large language models (LLMs) have achieved promising results on individual pa…
Large Language Models (LLMs) increasingly require selective removal of harmful or sensitiv…
Personalized interpretation of medical reports has emerged as an increasingly important ne…
Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard act…
Large language models often fail to answer questions about a bounded document collection w…
We present a novel approach to efficient LLM agent harness optimization through adaptive v…