ai.hackcv
论文精选 82arXiv

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models· 自动回归马赛克:测试纯文本模型的空间推理能力

Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or simply the ability to translate spatial descriptions into code. We introduce Autoregressive Mosaics (AM-Bench), a benchmark that separates these factors: First, a translation task gives a model a fully specified geometry of a picture in words as a prompt and asks for the code that produces it. Second, a layout task requires the model to compose an image from an underspecified prompt. Across eight open-weight text-and-code-only models, all models reliably translate specified geometry into code, but their open-ended layout performance differs substantially, indicating that thes

领域:cs.AI作者:Ashwin Nedungadi、Stefan Oehmcke、Stefan Lüdtke
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考