ai.hackcv
Agent / 智能体 · 自主智能体、工具调用与规划
论文精选 60已解读

A Deep Generative Model for Synthesizing Labeled Wireless Signals· 用于合成带标签无线信号的深度生成模型

Wireless signals with position-related labels are pivotal for both performance evaluation …

AI 解读一种生成带标签无线信号的新深度生成模型,适应不同环境,提升训练效果。
领域:cs.AI作者:Yuxiao Li、Keke Hu、Santiago Mazuelas
📎 arXiv🕒 09-05 01:44🔗 arxiv.org
论文精选 60已解读

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models· 在变压器语言模型中测量上下文 individuation 的工具箱技术手册

A transformer language model assigns a single, context-independent vector to a word type a…

AI 解读探索变压器语言模型中词的上下文特化表征工具箱
领域:cs.AI作者:José Luciano Verçosa Marques、Frederico Jorge Heitmann、Daniel Omar Perez
📎 arXiv🕒 09-05 00:38🔗 arxiv.org
论文精选 60已解读

Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness· 大型语言模型在建筑能源系统中的应用

Building automation systems generate rich sensor data yet remain insight-poor because hete…

AI 解读探讨大型语言模型在建筑能源系统中的应用现状与挑战。
领域:cs.AI作者:Alexander Neubauer、Tianzhen Hong、Han Li
📎 arXiv🕒 09-05 00:08🔗 arxiv.org
论文精选 65已解读

A Computationally Feasible Framework for Causal Probabilistic Explanation· 可计算的因果概率解释框架

Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, …

AI 解读提出了一种结合因果性和概率性的可计算解释框架PCI,以解决大规模模型的解释性问题。
领域:cs.AI作者:Rafal Urbaniak、Sam Witty、Daniel Waxman
📎 arXiv🕒 09-04 01:55🔗 arxiv.org
论文精选 65已解读

Efficient Test-Time Adaptation through Human-AI Interaction· 通过人机交互实现高效的测试时适应

AI agents are trained on population-scale data to encode broad capabilities spanning those…

AI 解读通过人机交互实现高效测试时适应,提升AI系统个性化表现
领域:cs.AI作者:Zora Zhiruo Wang、Apurva Gandhi、Rulin Shao
📎 arXiv🕒 09-04 01:33🔗 arxiv.org
论文精选 65已解读

Environment Evolution for Terminal Agents· 终端代理的环境进化

Scaling interactive and verifiable environments is critical for training terminal agents. …

AI 解读提出一种提升终端代理训练效果的环境进化方法。
领域:cs.AI作者:Zhiyuan Fan、Tinghao Yu、Yuanjun Cai
📎 arXiv🕒 09-04 01:26🔗 arxiv.org
论文精选 82

Spurious Advantage Hidden in GRPO· GRPO 中的虚假优势

Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning wit…

领域:cs.AI作者:Jiamian Wang、Samyadeep Basu、Koustava Goswami
📎 arXiv🕒 09-04 00:37🔗 arxiv.org
论文精选 65已解读

Discriminative World Models for Web Agents· 网络代理的世界模型

Recent web agents use world models for test-time action selection by sampling candidate ac…

AI 解读提出了网络代理世界模型新训练目标,提高状态辨识度。
领域:cs.AI作者:Kelvin Li、Dhruv Pendharkar、Anish Pahilajani
📎 arXiv🕒 09-03 01:59🔗 arxiv.org
论文精选 65已解读

AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application· AI上下文测量框架

Researchers increasingly use artificial intelligence to construct measures of social, orga…

AI 解读AI上下文测量框架评估AI衍生度量在情境模型中恢复个体和群体效应的能力。
领域:cs.AI作者:Wenxin Jiang、Xuyang Wang、Yuxiao Wu
📎 arXiv🕒 09-03 01:09🔗 arxiv.org
论文精选 65已解读

Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents· 用于工厂代理的测量驱动子网络选择

On-premise assistants can give factory workers conversational access to machine documentat…

AI 解读研究工厂代理助手的子网络选择方法,提升小型设备上的问答性能。
领域:cs.AI作者:Vasileios Rizeakos、Georgios Paisios、Alexandros Machairas
📎 arXiv🕒 09-03 00:00🔗 arxiv.org
论文精选 65已解读

Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting· Loom:通过嵌入空间重加权将诊断信息编织成自由文本共识

Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundament…

AI 解读Loom 通过嵌入空间重加权整合多个诊断模板的输出,形成可靠的自由文本共识。
领域:cs.AI作者:Ron Begleiter、Katya Egert Berg、Gilad Saban
📎 arXiv🕒 09-02 22:24🔗 arxiv.org
论文精选 65已解读

Collective creativity in hybrid societies· 混合社会中的集体创造力

Generative AI is changing how cultural artifacts are created and circulated, and with it o…

AI 解读探讨混合社会中AI对集体创造力的影响
领域:cs.AI作者:Mason Youngblood、Katie Mudd、Manuel Anglada-Tort
📎 arXiv🕒 09-02 21:59🔗 arxiv.org
论文精选 65已解读

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI· CivBench:文明VI中工具媒介代理的长期基准测试

We present CivBench, an open-source benchmark for evaluating language model agents in long…

AI 解读一个用于评估文明VI中语言模型代理长期表现的基准测试框架。
领域:cs.AI作者:Austin Tudor David Andrews、Liam Wilkinson、Jamie Heagerty
📎 arXiv🕒 09-02 19:31🔗 arxiv.org
论文精选 65已解读

UTP-Bench: Uncertainty-aware Travel Planning Benchmark· UTP-Bench: 面向不确定性的旅行规划基准

Large Language Models (LLMs) have recently demonstrated strong capabilities in automated t…

AI 解读提出面向不确定性的大规模旅行规划基准UTP-Bench
领域:cs.AI作者:Etcharla Revanth Rao、Priyanshu Karmakar、Shubhojit Mallick
📎 arXiv🕒 09-02 18:42🔗 arxiv.org