Compile by Training: Turning Natural-Language Specifications into Local Neural Functions· 编译时训练:将自然语言规范转换为本地神经函数
Many recurring text functions are easy to describe but difficult to implement with rules, …
Many recurring text functions are easy to describe but difficult to implement with rules, …
Language-model judges now gate training data, score generations, and drive leaderboards. T…
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appen…
Reasoning traces from chain-of-thought models appear to offer a legible window into how a …
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge dur…
Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, …
For scientific progress, we need benchmarks that test the limits of state-of-the-art model…
On-policy distillation (OPD) combines student-generated rollouts with dense token-level su…
Multi-agent AI science ecosystems rely on agents possessing tools that allow them to commu…
Research and news coverage of language-model deception increasingly attributes human-like …
As terminal-based code agents become prevalent, agent trajectories have accumulated at sca…
AI agents are trained on population-scale data to encode broad capabilities spanning those…
AI agents are increasingly being developed and deployed across organizations using heterog…
Scaling interactive and verifiable environments is critical for training terminal agents. …
Large language models are increasingly used to support organizational decisions, yet users…
Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GD…
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic c…
Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning wit…
IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domai…
Procedural instruction following is a basic requirement for controllable language-model sy…
Evaluating large language models (LLMs) in safety-critical, physics-governed environments …
For trained operators, gauge reading requires little specialized knowledge, low cognitive …
Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet m…
We present an alternative characterization of the occupancy measure of reinforcement learn…
An image editor may satisfy every regional plausibility constraint separately even when no…
Distill your knowledge, memories, and decisions into an open-source, inspectable AI Agent …
AI 测试开发学习路线大全:测试工程师的 AI 时代进阶指南,涵盖 AI 辅助测试、自动化测试、Agent 测试开发等,持续更新 🚀
Spark-x2.5 open model series. Pushing the Limits of Agentic Capabilities in On-Device Mode…
当AI开始“预演”一场暴雨
AI直接吐出正确答案,但最关键的可能不是答案
如此“反骨”的方法,具体又是怎么实现的?
最后靠Harness救回来
OpenAI’s latest agent swarm incident adds urgency to calls for independent investigations …
Nscale, which recently struck a $45 billion deal with Anthropic, is in talks to raise addi…
It's the latest failure of OpenAI's internal monitoring and security systems.
It’s officially the Ternus era at Apple. Tim Cook stepped down as CEO this week, handing t…
Introducing GPT-6 Astra, our most intelligent and aligned model yet, with state-of-the-art…

YouTube · OpenAI

YouTube · OpenAI