Learning When to Trust via Selective Context Preference Optimization· 学习何时信任:通过选择性上下文偏好优化
Language models increasingly condition their answers on external signals, and a single mis…
Language models increasingly condition their answers on external signals, and a single mis…
Tool use transforms LLMs into agents that act beyond their training data, and for code-cap…
Electronic health record (EHR) feature engineering is a major bottleneck in clinical resea…
Real-world video benchmarks provide broad coverage, but their fixed clips entangle event c…
This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with res…
Multilingual reasoning transfer is crucial for extending reasoning capabilities of large l…
LLM-based agentic systems have shown remarkable capabilities in complex domains, while suf…
Task-oriented conversational agents are evaluated using curated or automatically generated…
Large language models (LLMs) increasingly support complex professional tasks, yet their ca…
Retrieval-augmented generation over long documents is dominated by one design: chunk the t…
As LLMs are increasingly deployed within agentic systems, their capabilities depend not on…
Automatic speaking assessment systems are increasingly deployed in high-stakes settings to…
Cardiac arrest remains one of the most lethal conditions encountered in intensive care uni…
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations s…
Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks an…
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities …
While deep learning models, particularly transformer-based architectures, have shown impre…
Training large language model agents for long-horizon tool use typically relies on interac…
Agents backed by large skill libraries must decide which skills to load and in what order.…
Microarchitecture design space exploration suffers from expansive search spaces and expens…
We present a schema-based framework for extracting complex, structured information from un…
Synthetic 3D scene generation is increasingly used as a data source for computer vision an…
Earth-surface monitoring requires change detection models capable of recognizing arbitrary…
End-to-end document parsers provide a unified interface, but serialize page layouts and re…
Most agent benchmarks evaluate tasks independently and cannot measure whether experience f…
Graph-Orchestrated Agent Loop — a production-grade framework on LangGraph. Combine workflo…
An agentic LLM-powered knowledge assistant that enhances RAG capabilities through automate…
海外团队纷纷下注,NeoLab浪潮狂飙
「我,Jeff Dean,打钱」
唉,赶紧整完快点发布吧。。。
再花15亿美元买现成AI编程团队
啊,这真是个AI“失控”的夏天
OpenAI said it has suspended work on some aspects of its upcoming model Astra over concern…
After its own AI usage wake-up call, Rippling this week unveiled AI Spend Console, a produ…
Cloudflare has introduced Kitesurf, a cloud-hosted browser designed for AI agents instead …
<img data-src="https://static.leiphone.com/uploads/new/images/20260807/6a75f97f62ce2.jpg" …
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taki…
Historian Jill Lepore has a theory about why tech companies often use soaring language to …

YouTube · NVIDIA

YouTube · NVIDIA

YouTube · NVIDIA