CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes· CritICL:小模型失败模式改进大模型推理
Recent advances in inference-time scaling have significantly improved the reasoning perfor…
Recent advances in inference-time scaling have significantly improved the reasoning perfor…
Agent skills package specialized knowledge and workflows into reusable resources that exte…
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy …
Chemical reactions are fundamentally transformations in electron space, yet most machine l…
Transduced language models (TLMs) compose a pretrained \emph{source} language model with a…
Currently used sepsis severity indices rely on fixed variables and weights established dec…
Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning…
Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of la…
LLMs are increasingly able to answer complex questions about enterprise-scale document col…
We study how sophistication in generative AI (genAI) use varies among the back-office work…
Steering interventions targeting eval-awareness, a model's recognition that it is being te…
Agent harnesses shape how language-model agents use instructions, tools, and runtime compo…
We ask whether large language models (LLMs) can design effective algorithms for well-speci…
Although Large language models (LLMs) mediate access to knowledge and computational assist…
Efficiently improving autonomous agents across diverse tasks is central to accelerating re…
LLM agents increasingly rely on generated interaction data to learn how to interact with e…
An LLM agent shown a professional-looking market panel commits to a directional call on a …
Conversational AI systems, such as chatbots and virtual assistants, are becoming increasin…
The development of frontier models is commonly perceived to be the exclusive remit of a sm…
Tool-augmented LLM agents must rely on untrusted runtime Observations to complete open-end…
In recent years, graph anomaly detection (GAD) based on frequency-domain filtering have ac…
Despite their potential in standardized graph tasks, Large Language Models (LLMs) remain b…
Internet memes are a pervasive form of multimodal online communication; however, such comm…
Large Language Models (LLMs) operate in hospitals, courtrooms, banks, and public service d…
Ontology learning from text remains challenging despite significant progress in Large Lang…
Open-source multi-provider AI gateway for Claude Code and other coding agents, with model …
Open-source AI-assisted per-application routing for Windows with OpenAI/Ollama drafts and …
Unbagrnd - Free, Fast & Open-source AI-powered background remover for Windows, macOS, and …
AI Chat with powerful models for free
Emotion Ball 是一套面向 AI 助手的表情引擎:32 种状态表情全部由纯 SVG 与原生 JavaScript 实时驱动,零框架、零图片资源。AI 侧只需输出一个 em…
推理软件栈的微小差异,就能改变输出token
本科读车辆工程,如今教AI第一班
The new generation of data center systems is increasing efficiency with smarter traffic co…
AI「自进化」,越来越近了
Coding正在变成Al世界的数字执行力
从剧本到成片
华为终端近日正式宣布,HarmonyOS 7|HUAWEI Mate XT 2 及全场景新品发布会将于9月7日举行,引发大众对全新三折叠的广泛关注与讨论。随后,8月29日核心商圈户…
8月28日,在2026中国国际大数据产业博览会期间,以“从数据到Agent,数据价值链的重构与创新”为主题的数据要素大会在贵阳举办。政府机构、头部大数据集团、行业企业及专家学者齐聚…
Neocloud Lambda has raised $1B in private debt to buy Nvidia AI chips and lease them to Mi…
Given 10 benchmarks for specific misaligned behaviors, the automated systems were able to …
There's a lot of capital pouring into the business of giving models away.
8月28日,佑驾创新(2431.HK)发布2026年度中期业绩报告,交出营收毛利双增、亏损收窄、无人车收入爆发的中期答卷,规模效应加速释放,盈利质量与经营韧性同步提升。面对汽车智能…
Our decision to wind down our contract providing OpenAI models to Cursor following its acq…

YouTube · OpenAI

YouTube · OpenAI

YouTube · OpenAI

YouTube · OpenAI

YouTube · OpenAI

YouTube · OpenAI