When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models· 当验证的世界模型仍失败
Large language models can synthesize a game's rules as executable code - a Code World Mode…
Large language models can synthesize a game's rules as executable code - a Code World Mode…
Bayesian Belief Networks (BBNs) are powerful tools for decision-making under uncertainty. …
Financial sentiment extraction has largely relied on news text and supervised extraction a…
This paper presents a novel three level hierarchical learning architecture for autonomous …
World models are widely used in offline reinforcement learning (RL) to improve sample effi…
Type 1 Diabetes (T1D) is a chronic, life-threatening autoimmune condition characterized by…
Feedback-driven loops support iterative improvement in large language models, reinforcemen…
Although large language models (LLMs) have set benchmarks for zero-shot reasoning, their d…
Retrieval over corpora that mix several domains often returns relevant but wrong-domain ev…
An agent harness is the external control layer that turns a base LLM into an executable ag…
Reliable, low-latency uplink connectivity is a key requirement for C-V2X networks in dense…
Reinforcement learning has emerged as the dominant paradigm for training large language mo…
Wildfire detection from satellite imagery is a semantic image segmentation problem that ha…
Tool-augmented large language model agents excel at long-horizon tasks, yet they are typic…
Inspired by how humans communicate spatial information, language-guided geo-localization h…
In predictive modeling, the ability to explain why a model produces a given target predict…
Retrieval Augmented Generation (RAG) has proven to be a widely successful process at impro…
Representative clutter height (RCH) is a key parameter in radio propagation and interferen…
Pre-trained vision-language models (VLMs) enable zero-shot image classification by computi…
Despite the proliferation of Explainable AI (XAI) techniques -- from feature attributions …
Embodied cognition requires agents to connect high-level task reasoning with the physical …
Recent advances in Large Language Models have fueled autonomous AI agents capable of tackl…
This position paper explores how Agentic AI and Model Context Protocol (MCP) can support p…
The Platonic Representation Hypothesis (PRH) holds that as models scale, representations o…
We introduce RegNetAgents, an AI-oriented multi-agent framework for structured, query-driv…
Large language models (LLMs) increasingly serve as high-level planners for embodied agents…
Annotation quality is a major bottleneck in building reliable and explainable artificial i…
Graphical user interface (GUI) automation remains challenging in real-world environments, …
AI benchmarks increasingly leverage item-level statistical models, particularly item respo…
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations…
Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench …
Artificial intelligence is transforming scientific research - not merely as a more powerfu…
World models are usually evaluated as components of model-based reinforcement learning (MB…
Parameter-efficient fine-tuning reduces model and optimizer memory, but dense attention st…
Understanding the brain increasingly depends on integrating evidence across scales, modali…
The integration of AI-driven systems in creative work has sparked debates among artists an…
The deployment of autonomous cyber-physical systems in safety-critical environments requir…
This paper suggests the adoption of a novel inversion in AI ethics: instead of asking how …
Per-subgroup fairness audits of medical image classifiers face a sample-size problem: mino…
Channel foundation models (CFMs) are developing rapidly, with recent studies reporting ben…
Automated optimisation is increasingly adopted in industrial processes, yet a trust gap pe…
Open-source AI coding agent for the terminal. Claude Code-grade accuracy with smart model …
The world's #1 vibe coding benchmark — models go head-to-head on real engineering tasks, j…
Current AI, a non-profit building AI that leaves no one culture behind, has made remarkabl…
智能进了社会,治理不能慢半拍
主观世界模型
奇异摩尔首次亮相WAIC 2026
2026年7月17日,中国上海——2026世界人工智能大会暨人工智能全球治理高级别会议(WAIC 2026)在沪开幕。燧原科技以“芯火燎原”为主题参展WAIC 2026(世博展览馆…
随着科学问题日益复杂、科研数据快速增长,传统科研方式面临效率瓶颈,AI for Science正推动科学研究从经验驱动走向数据与模型驱动。 7月18日,在WAIC2026中科闻歌决…
还得是AI圈春晚
Om AI联汇倡议物理AI协同发展,端侧架构受重视。
2026世界人工智能大会专题论坛聚焦五大议题。
在2026世界人工智能大会(WAIC 2026)上,安顿健康以两项首发引发广泛关注:我国首个生命预警表标准,以及行业首个七诊合参中医机器人——安顿中医机器人。 WAIC现场图片 两…
极响应全国一体化算力网建设部署
《关于成立世界人工智能合作组织的协定》签署仪式16日在上海举行,来自亚洲、非洲、拉美和欧洲的29个国家现场签署协定,成为创始成员国。记者从有关方面获悉,这标志着中方倡导推动成立世界…
36氪获悉,新易盛公告称,预计2026年半年度归属于上市公司股东的净利润为70.00亿元-80.00亿元,同比增长77.56%-102.93%。业绩变动主要系受益于人工智能相关算力…
36氪获悉,7月19日,阿里最新一代大模型千问3.8即将发布并开源。目前,Qwen3.8-Max预览版已率先上线阿里云Token Plan、Qoder及QoderWork,用户可提…
黎曼动力正式发布Riemann-1.0
36氪获悉,近日,2026世界人工智能大会(WAIC2026)期间,中科闻歌发布“基、枢、核、脑、端”完整AI决策产品体系。该体系以DOMA架构为技术底座,覆盖数据治理、业务建模、…
在2026世界人工智能大会分论坛“长三角人工智能产业集群发展论坛”上,长三角人工智能协同投资平台共建签约,签约方包括长三角投资公司、国家开发投资集团、上海国投公司、江苏省高投集团、…
7月19日,2026年世界人工智能大会正在进行中,高通公司全球副总裁、中国区研发负责人、IEEE Fellow徐晧博士出席“端侧AI创新和行业发展论坛”,并发表《智能体AI时代:计…
Harness本身也可以被搜索、验证和迭代
端到端整合、能跨场景复用的操作方法论
星宸科技公告称,预计2026年半年度归属于上市公司股东的净利润为8.20亿元-9.00亿元,同比增长583.72%-650.42%。业绩变动主要系端边侧AI产业需求全面爆发,公司产…
中科创达公告称,预计2026年半年度归属于上市公司股东的净利润为2.38亿元-2.46亿元,同比增长50%-55%。业绩变动主要系公司坚持全球化及AIOS战略,积极拓展全球业务,深…
AI技术在生物分子设计中的应用进展。
工业AI应用深化,从单一场景到全面生产力提升。
36氪获悉,7月18日,商汤科技正式发布旗舰级大模型SenseNova U1 Pro(以下简称U1 Pro)。U1 Pro定位为面向长程任务的交付级原生多模态智能体基座,实现理解、…
摩尔线程展示AI训练推理成果,强调算力价值。
<img class="rich_pages wxw-img" data-aistatus="1" data-backh="300" data-backw="546" data-c…
STEPX Neo亮相,展示先进人机交互技术。
36氪获悉,7月18日晚间,一目科技宣布完成超10亿元E轮融资,估值突破100亿元。据悉,本轮融资由多家一线人民币基金、头部美元基金和知名产业方联合参与,所融资金将用于具身智能触觉…
7月19日,2026世界人工智能大会——地震人工智能论坛在上海召开。会议聚焦人工智能与防震减灾深度融合发展,集中展示了人工智能技术在防震减灾领域的最新科研成果和实际应用。中国地震局…
36氪获悉,7月19日,2026世界人工智能大会期间,蚂蚁集团首席执行官韩歆毅围绕企业AI如何从第一轮部署走向深度价值落地,分享了三点思考:企业AI的智能不止于知识累积,更在于经验…
7月18日,在2026世界人工智能大会上,上海标志性太空数字基建工程“星枢计划”首发星座正式对外发布。仪式由WAIC组委会与复旦大学联合主办,作为上海重点布局的太空算力产业标杆项目…
7月19日,阿里云百炼正式上线可实时构建和交互的开放式世界模型产品HappyOyster 1.0(快乐生蚝1.0)。该模型深度学习物理世界状态转移规律,能主动推演从动作到反馈的因果…
36氪获悉,7月17日,WAIC世界人工智能大会上,百度一镜现场首发数字人视频播客解决方案,突破数字人微表情技术,呈现自然插话、专业影视视听语言能力。该方案解决了真人访谈成本高、传…
7月19日,在2026世界人工智能大会上,蚂蚁数科发布新一代企业 AI 解决方案——智能体超级工厂核心平台 Agentar 2.0。据悉,Agentar2.0首批预置近200个岗位…
WAIC现场叠箱封胶,解锁物理AI新技能
2026世界人工智能大会上,各类人工智能终端令人目不暇接,但记者注意到,行业关注的焦点已经从人工智能技术的展示转向相关技术如何更好地应用在生产和生活中。专家指出,数据规模、评测标准…
理解、生成、行动的原生统一
一家CPU芯片公司,为什么要做操作系统? 7月17日,此芯科技在WAIC 2026首日发布了AGX Agentic Compute战略,并推出AGX Station等产品。比起一台…
PuduFM+PuduAgent,一并在不同本体上持续落地,共同构成了普渡机器人的顶层战略「一脑多形」。
据阶跃官微消息,7月18日,阶跃星辰与上海期智研究院宣布正式达成合作,双方将共同设立智能体前沿研究院,围绕智能体网络及经济原理、AI Safety等方向开展联合研究,探索Agent…
Chinese company Moonshot AI released a new version of its Kimi model this week, prompting …
7月18日,云天励飞在2026 WAIC公布未来两年多AI推理芯片路线图。公司计划推出DeepVerse100P、DeepVerse100D、DeepVerse100L三款芯片,分…
机器人真正具备了干活的完整能力