ai.hackcv
AI 解读

由 LLM 对每日入选的 AI 论文与开源项目做结构化中文解读,点卡片看完整内容。

论文已解读

A Deep Generative Model for Synthesizing Labeled Wireless Signals· 用于合成带标签无线信号的深度生成模型

一种生成带标签无线信号的新深度生成模型,适应不同环境,提升训练效果。

📎 arXiv🕒 9月4日
论文已解读

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe· 韩国开放公共API的多步调用基准

韩国开放公共API的多步调用基准及其数据生成方法

📎 arXiv🕒 9月4日
论文已解读

Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence· 必要还是充分?用行为证据评估LLM解释

评估LLM解释的必要性和充分性

📎 arXiv🕒 9月4日
论文已解读

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models· 分子属性预测中的大模型文字级检索

大模型在分子属性预测中普遍存在文字级数值检索现象

📎 arXiv🕒 9月4日
论文已解读

CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents· CUA-Universe: GUI+CLI 混合代理的可扩展动态环境

提出 CUA-Universe, 用于 GUI+CLI 混合代理的可扩展动态环境

📎 arXiv🕒 9月4日
论文已解读

Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education· 谁来评价我的作业?学生对透明AI写作评估的观点

本研究探讨了学生对AI写作评估系统的接受度和理解情况。

📎 arXiv🕒 9月4日
论文已解读

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability· 代理的记忆能否在模型升级中幸存?

研究模型升级对代理记忆的影响,探索不同记忆存储方式的迁移性。

📎 arXiv🕒 9月4日
论文已解读

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models· 在变压器语言模型中测量上下文 individuation 的工具箱技术手册

探索变压器语言模型中词的上下文特化表征工具箱

📎 arXiv🕒 9月4日
论文已解读

LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams

利用大模型生成启发式用于量子电路合成,优化可逆门电路的成本。

📎 arXiv🕒 9月4日
论文已解读

Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness· 大型语言模型在建筑能源系统中的应用

探讨大型语言模型在建筑能源系统中的应用现状与挑战。

📎 arXiv🕒 9月4日
论文已解读

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation· RISE:通过自外推策略蒸馏递归改进

通过自外推策略蒸馏实现语言模型的递归改进

📎 arXiv🕒 9月4日
论文已解读

TokenMatch: 3D Mesh Correspondence Transformer with Curvature-Guided Tokenisation· TokenMatch:曲率引导标记化的3D网格匹配Transformer

提出TokenMatch,通过曲率引导标记化解决3D网格匹配问题。

📎 arXiv🕒 9月3日
论文已解读

Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision· 时间自蒸馏:无监督视频视觉状态跟踪

提出S$^3$T,利用时间自蒸馏无监督学习视频视觉状态跟踪。

📎 arXiv🕒 9月3日
论文已解读

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction· Scal3R: 提升在线 3D 重建效率

提出 Scal3R 以多参考相对姿态查询方式提升在线 3D 重建效率。

📎 arXiv🕒 9月3日
论文已解读

Principia: Relational Physics Tests for Video Models· Principia:视频模型的物理测试

视频模型物理推理测试的新方法与基准

📎 arXiv🕒 9月3日
论文已解读

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints· 清洁工程,不稳定测量:黑盒LLM观察者可靠性测试失败

黑盒LLM观察者在共享端点上的可靠性测试失败。

📎 arXiv🕒 9月3日
论文已解读

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States· Puffin-World:用原生3D世界状态扩展统一多模态模型

Puffin-World 提出了一种集成了物理理解、空间模拟和3D世界生成的多模态模型。

📎 arXiv🕒 9月3日
论文已解读

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing· 一个编辑器,多种编辑:多样化视频编辑的统一无训练框架

多种视频编辑任务无训练统一框架

📎 arXiv🕒 9月3日
论文已解读

Robust PAC Learning of Concurrent Stochastic Games· 并发随机游戏的鲁棒 PAC 学习

并发随机游戏的鲁棒 PAC 学习框架,解决纳什均衡存在性问题。

📎 arXiv🕒 9月3日
论文已解读

Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning· 在弱监督密集视频标注中以VLM引导过渡事件发现

提出SBS框架改善视频事件描述与定位精度

📎 arXiv🕒 9月3日
论文已解读

A Computationally Feasible Framework for Causal Probabilistic Explanation· 可计算的因果概率解释框架

提出了一种结合因果性和概率性的可计算解释框架PCI,以解决大规模模型的解释性问题。

📎 arXiv🕒 9月3日
论文已解读

Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations· 使用3D基础模型进行零样本新视角深度合成

利用3D基础模型的内部表示进行新视角的零样本深度预测。

📎 arXiv🕒 9月3日
论文已解读

Rethinking On-Policy Distillation of Large Language Models II: One Training Example· 重新思考大语言模型的单样本策略蒸馏

单样本策略蒸馏在大语言模型训练中表现出显著效果。

📎 arXiv🕒 9月3日
论文已解读

A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms· 多智能体自主研究群体中的作弊与举报

探讨多智能体自主研究群体中的作弊与举报行为。

📎 arXiv🕒 9月3日
论文已解读

Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs· Para-Pipe:利用 SoC 上 ML 计算图的层次并行性

通过层次并行性优化 SoC 上的 ML 计算图性能。

📎 arXiv🕒 9月3日
论文已解读

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents· SWE-Gate: 代码代理仅通过功能测试不够

功能测试不足以评估代码代理的全面能力。

📎 arXiv🕒 9月3日
论文已解读

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

通过因果框架区分语言模型欺骗行为和机制

📎 arXiv🕒 9月3日
论文已解读

Parameterised graph theory for tensor networks: entanglement rerouting, structural simplification, and agnostic tomography· 参数化图论与张量网络

参数化图论在张量网络中的应用,优化状态表示与学习复杂度

📎 arXiv🕒 9月3日
论文已解读

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments· 终端宇宙:将代理轨迹转换为可扩展的终端环境

将代理的终端操作轨迹转换为可重用的执行环境和新任务。

📎 arXiv🕒 9月3日
论文已解读

Efficient Test-Time Adaptation through Human-AI Interaction· 通过人机交互实现高效的测试时适应

通过人机交互实现高效测试时适应,提升AI系统个性化表现

📎 arXiv🕒 9月3日
论文已解读

The Natural Language Interaction Protocol and Standard for AI Agents· AI 代理的自然语言交互协议与标准

提出 AI 代理通用自然语言交互协议

📎 arXiv🕒 9月3日
论文已解读

Environment Evolution for Terminal Agents· 终端代理的环境进化

提出一种提升终端代理训练效果的环境进化方法。

📎 arXiv🕒 9月3日
论文已解读

Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable· 大模型推荐的理论依据

大模型推荐如何在缺乏地面真实数据时评估可靠性。

📎 arXiv🕒 9月3日
论文已解读

Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM· 为什么Gated DeltaNet能承受4位量化

探索GDN模型在4位量化下的表现与优势

📎 arXiv🕒 9月3日
论文已解读

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training· 动态评分规则在长期代理训练中的应用

动态评分规则应用于长期代理训练,提高任务执行的细粒度归因。

📎 arXiv🕒 9月3日
论文已解读

Discriminative World Models for Web Agents· 网络代理的世界模型

提出了网络代理世界模型新训练目标,提高状态辨识度。

📎 arXiv🕒 9月2日
论文已解读

AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application· AI上下文测量框架

AI上下文测量框架评估AI衍生度量在情境模型中恢复个体和群体效应的能力。

📎 arXiv🕒 9月2日
论文已解读

Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis· LLM 用于电信根因分析

使用 LLM 优化电信网络根因分析,提高诊断准确性。

📎 arXiv🕒 9月2日
论文已解读

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment· SafeEvolve: 基于代理经验的安全对齐

基于代理经验的运行时控制与内在安全共进化框架

📎 arXiv🕒 9月2日
论文已解读

Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents· 用于工厂代理的测量驱动子网络选择

研究工厂代理助手的子网络选择方法,提升小型设备上的问答性能。

📎 arXiv🕒 9月2日
论文已解读

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems· 多智能体LLM系统的双层协调反射

探讨多智能体LLM系统中的协调机制与优化方法。

📎 arXiv🕒 9月2日
论文已解读

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills· GitHub 仓库到 AI 技能

将 GitHub 仓库知识提炼为 AI 研究技能。

📎 arXiv🕒 9月2日
论文已解读

Door-in-the-Face Requests and Refusal Behaviour in Large Language Models· 大型语言模型中的门面拒绝行为

研究大型语言模型中门面拒绝技巧的效果。

📎 arXiv🕒 9月2日
论文已解读

Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting· Loom:通过嵌入空间重加权将诊断信息编织成自由文本共识

Loom 通过嵌入空间重加权整合多个诊断模板的输出,形成可靠的自由文本共识。

📎 arXiv🕒 9月2日
论文已解读

Collective creativity in hybrid societies· 混合社会中的集体创造力

探讨混合社会中AI对集体创造力的影响

📎 arXiv🕒 9月2日
论文已解读

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI· CivBench:文明VI中工具媒介代理的长期基准测试

一个用于评估文明VI中语言模型代理长期表现的基准测试框架。

📎 arXiv🕒 9月2日
论文已解读

UTP-Bench: Uncertainty-aware Travel Planning Benchmark· UTP-Bench: 面向不确定性的旅行规划基准

提出面向不确定性的大规模旅行规划基准UTP-Bench

📎 arXiv🕒 9月2日
论文已解读

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers· 通过熵选择性代理引导:从未完美的 VLM 教师学习自主策略

通过选择性地利用视觉语言模型的建议来训练轻量级强化学习策略。

📎 arXiv🕒 9月1日
论文已解读

Can LLMs Discover Scientific Laws in Real and Parallel Worlds?· 大型模型能否发现科学定律?

LLMs能否发现科学定律及评估方法研究

📎 arXiv🕒 9月1日
论文已解读

EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation· EvoSCM:通过因果模型进化和实验的科学信念修订

AI科学代理通过因果模型的进化与实验学习科学信念修订。

📎 arXiv🕒 9月1日
论文已解读

When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation· 当防护栏看起来有效时:LLM 代理商业评估中的构念效度失败

LLM代理商业评估中,构念效度失败导致防护措施效果被高估。

📎 arXiv🕒 9月1日
论文已解读

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement· 多日自主软件开发的持续改进框架

多日自主软件开发中持续改进的框架

📎 arXiv🕒 9月1日
论文已解读

Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers· 解析流:长视界代理及其观察者的实时追踪模型

提出一种实时追踪模型,优化长视界代理的可观测性和资源效率。

📎 arXiv🕒 9月1日
论文已解读

EdiTikZ: Scientific Figure Editing from Revision Trajectories

利用修订轨迹生成大规模科学图表编辑数据集

📎 arXiv🕒 9月1日
论文已解读

Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations· 神经符号几何抽象:从观察到数学符号表达

神经符号几何抽象框架将几何观察转化为数学符号表示。

📎 arXiv🕒 9月1日
论文已解读

EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems· EDGE:基于错误依赖图的多错误归因

基于错误依赖图归因多智能体系统中的多个相关错误

📎 arXiv🕒 9月1日
论文已解读

SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding· SymFold: 结合进化和结构先验的蛋白质逆折叠

结合进化和结构先验,提出SymFold模型改善蛋白质逆折叠预测。

📎 arXiv🕒 9月1日
论文已解读

Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades· 廉价验证者,大盲点

研究了通过 cheap student 和 verifier 之间的闭合循环降低成本的可靠性和代价。

📎 arXiv🕒 9月1日
论文已解读

LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting· LEAP:基于LLM的概率预测的似然性提取与聚合

提出LEAP方法,改进LLM预测系统,通过独立检查证据项并聚合似然性来提高预测透明度和准确性。

📎 arXiv🕒 9月1日
论文已解读

OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques· OntoAligner-Ensemble: 跨异构本体对齐技术的投票融合

本体对齐技术的投票融合框架

📎 arXiv🕒 8月31日
论文已解读

When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning· 大模型规模何时有益?本体学习的控制研究

研究大模型规模对本体学习的影响

📎 arXiv🕒 8月31日
论文已解读

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing· BLOOM-WILT:大模型审计中的逻辑倾斜技术

提高大语言模型审计效率的技术,无需额外训练成本。

📎 arXiv🕒 8月31日
论文已解读

Cross-Regional Grapevine Cold Hardiness Prediction via Learned Multimodal Latent Representations· 跨区域葡萄藤抗寒性预测

利用学习的多模态潜在表示预测葡萄藤的抗寒性,实现跨区域和品种的模型迁移。

📎 arXiv🕒 8月31日
论文已解读

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data· 通过自适应结构化提高数据推理效率

通过自适应结构化提高数据推理效率,降低 LLM 代理成本。

📎 arXiv🕒 8月31日
论文已解读

Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization· 代理策略优化中过程监督与基于结果的信用的协调

通过调整过程监督与结果信用,提高代理策略优化的精度。

📎 arXiv🕒 8月31日
论文已解读

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence· 大型推理模型超越人类监督:迈向超智能的路径

研究大型推理模型在减少人类监督情况下的自适应学习。

📎 arXiv🕒 8月31日
论文已解读

Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores· 错误预测,正确答案:从崩溃的 LLM 序列分数中恢复证据

从大模型的失败中恢复正确的推理证据。

📎 arXiv🕒 8月31日
论文已解读

Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents· 先测量再管理:编码代理工作内存评估

研究编码代理工作内存的语义异质性及管理策略。

📎 arXiv🕒 8月31日
论文已解读

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents· MNIST-PRO:AI代理的部分可观测世界

MNIST-PRO将经典的MNIST数据集转化为部分可观测的环境,以评估AI代理的感知与记忆能力。

📎 arXiv🕒 8月31日
论文已解读

Learning Action Models with Conditional and Quantified Effects via Uncertainty-Guided Exploration· 通过不确定性引导探索学习条件和量化效果的动作模型

通过不确定性引导探索,改进条件和量化效果的动作模型学习方法。

📎 arXiv🕒 8月31日
论文已解读

CARVE: Verified Expansion for Variable-Length Generation in Diffusion Language Models· CARVE:扩散语言模型中的可验证变长生成

扩散语言模型生成长度可调的新方法。

📎 arXiv🕒 8月31日
论文已解读

Logos: An Agent Harness on a Cross-Process Bus· Logos: 跨进程总线上的代理绑定

跨进程总线上的代理绑定机制,实现更灵活的代理系统。

📎 arXiv🕒 8月28日
论文已解读

InstructMesh: Selective Refinement of Generative 3D Models for Fabrication· InstructMesh:生成 3D 模型的交互式精炼工具

用于生成 3D 模型后制作精炼的交互工具。

📎 arXiv🕒 8月28日
论文已解读

When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI· 当机器人听错我们:语音控制具身 AI 的安全风险

研究语音识别错误对具身AI模型安全性的影响。

📎 arXiv🕒 8月28日
论文已解读

Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration· 提高MoE模型通信效率的训练方法

研究提高MoE模型通信效率的训练方法

📎 arXiv🕒 8月28日
论文已解读

AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction· AcrossVAM1.0: 文本辅助机器人视频预测的粒子世界建模

文本辅助分解视频预测任务,提高机器人预测视频精度与细节。

📎 arXiv🕒 8月28日
论文已解读

COVER: Identifiable Evaluation of Coalition Routing· 联合路由的可识别评估

提出联合路由可识别评估方法,提高多智能体系统性能评估准确性。

📎 arXiv🕒 8月28日
论文已解读

Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning· 学习使用工具:强化学习以实现工具集成的数学推理

强化学习结合工具使用,提升数学推理能力

📎 arXiv🕒 8月28日
论文已解读

Prove2Me: An Open Collaborative Platform for Scaling Math Formalization· Prove2Me:开放协作的数学形式化平台

开放协作平台 Prove2Me 促进数学形式化,降低门槛,实现人机协作。

📎 arXiv🕒 8月28日
论文已解读

Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs· 带可验证奖励的程序学习:符号反向传播

提出 PLVR,通过符号反向传播从输入输出示例中学习可验证程序。

📎 arXiv🕒 8月28日
论文已解读

VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings· VERA-8B:基于 SEC 文件的证据支持审计风险推理

VERA-8B 通过结合 SFT 和 GRPO 方法,基于 SEC 文件提供证据支持的审计风险预测。

📎 arXiv🕒 8月28日
论文已解读

RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents· RetailAgent:自我条件多模态 LLM 交易代理的结构化不利时机

研究 LLM 交易代理在金融市场中的预测性和不利时机。

📎 arXiv🕒 8月28日
论文已解读

Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation· 电商大规模杂货推荐的时序重购预测

使用生存模型预测大规模电商平台上杂货的时序重购行为。

📎 arXiv🕒 8月28日
论文已解读

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution· WikiSkill:将代理经验编译为持续知识

将代理经验整合为持久知识,促进技能进化。

📎 arXiv🕒 8月27日
论文已解读

Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation· 通过离散流匹配预测化学反应机制

通过电子占据空间的离散流匹配预测化学反应机制。

📎 arXiv🕒 8月27日
论文已解读

Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study· 无需逐小时监督的连续败血症严重程度评分学习

提出无需逐小时监督的败血症严重程度评分模型。

📎 arXiv🕒 8月27日
论文已解读

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases· 企业级 Q&A 基准测试

评估大规模企业文档集的LLM Q&A性能的新基准。

📎 arXiv🕒 8月27日
论文已解读

Sophistication in GenAI Use: Field Evidence from a Large Firm· 大公司生成式AI使用成熟度:实证研究

大型公司后办公室员工使用生成式AI的成熟度差异研究。

📎 arXiv🕒 8月27日
论文已解读

Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance· 并非所有评估意识都相同:能力框架预测合规性

评估意识的能力框架预测合规性,与安全框架预测不同。

📎 arXiv🕒 8月27日
论文已解读

Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification· 更智能验证,更进一步进化

通过行为感知验证优化代理框架,提高性能并减少资源浪费。

📎 arXiv🕒 8月27日
论文已解读

LLMs Can Design Near-Optimal OR Algorithms· 大模型能设计近优运筹算法

大模型能设计接近最优的运筹算法

📎 arXiv🕒 8月27日
论文已解读

BrailleBench: Investigating Multi-Criteria Braille Comprehension in Large Language Models· 盲文bench:多标准盲文理解评估

评估大语言模型对盲文的理解能力

📎 arXiv🕒 8月27日
论文已解读

Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search· 朴素提示优化:重新思考复杂提示搜索的必要性

一种简单的提示优化方法,不依赖复杂搜索,提高AI自主性能。

📎 arXiv🕒 8月27日
论文已解读

What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents· 什么是好的代理数据?

提出代理数据的两层框架与生成范式,提升 LLM 代理交互数据质量。

📎 arXiv🕒 8月27日
论文已解读

Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable· LLM在虚构证据面前的预测倾向

LLM在面对虚构证据时更倾向于做出无法证实的预测。

📎 arXiv🕒 8月27日
论文已解读

Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings· 行星预测引擎:通过智能数据选择和基础模型嵌入自动进行地理空间预测

利用智能数据选择和基础模型嵌入实现自然语言查询驱动的自动地理空间预测。

📎 arXiv🕒 8月26日
论文已解读

SwarmWorld: Stigmergic technological evolution in societies of language-model agents· SwarmWorld:语言模型代理的社会技术演化

研究基于语言模型的代理在共享环境中自我组织形成技术社会的能力。

📎 arXiv🕒 8月26日
论文已解读

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems· LLM数据代理的Trace Integrity

LLM数据代理的可靠性评估新标准:Trace Integrity。

📎 arXiv🕒 8月26日
论文已解读

Imitation Learning for Connection-Tableau Construction· 连接表构建的模仿学习

使用图神经网络和模仿学习优化定理证明的连接表构建政策。

📎 arXiv🕒 8月26日
论文已解读

AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs· AsymSpec:面向代理的LLM的上下文不对称投机解码

代理人LLM的上下文不对称投机解码框架,降低推理成本同时保持任务准确性。

📎 arXiv🕒 8月26日