世界模型碎片化:通用智能体的结构性认证方法

摘要
在“大世界”场景下,智能体无法具备通用能力,其能力必然专精于碎片化的世界模型。标准均匀保证无法区分关键瓶颈与无关失败。本文通过证明通用智能体并非万能,引入结构性认证框架,将有界目标条件性能映射到智能体内部世界模型的逐项保证,并给出算法实现。
背景解释
该研究针对通用智能体在复杂环境中的可靠性问题,提出结构性认证方法,通过局部化关键转换来确保长期规划的可信度。这对AI安全部署具有理论指导意义,尤其适用于需要高可靠性的自主系统。
原文译文
以下为抓取到的原文内容译文,已统一为站内阅读格式。
arXiv is now an independent nonprofit!Learn more×
Search arXiv
Press Enter to search ·Advanced search
Computer Science > Artificial Intelligence
arXiv:2606.24842v1 (cs)
[Submitted on 23 Jun 2026]
Title:World Models in Pieces: Structural Certification for General Agents
Authors:Yikai Lu,Yifei Wu,Xinyu Lu,Tongxin Li
View a PDF of the paper titled World Models in Pieces: Structural Certification for General Agents, by Yikai Lu and Yifei Wu and Xinyu Lu and Tongxin Li
Abstract:In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish between the understanding of critical bottlenecks and irrelevant failures. We first formalize this limitation by proving that general agents are not universal, rendering standard worst-case analysis uninformative. To overcome this, we introduce structural certification, a transition-local framework that maps bounded goal-conditioned performance to entry-wise guarantees on the agent's internal world model. Our main contribution is constructive. We provide algorithms that filter specific transitions using deep compositional goals and prove that a general agent on these goals has a structural world model with a O(1/n)+O(δ) error bound. Conversely, this bound is tight in the small-δ regime, whose existence is explicitly guaranteed by our certification. These results enable the certifiable deployment of general agents by localizing the specific transitions where long-horizon planning is reliable.
| | | | --- | --- | | Comments: | 30 pages, camera-ready version in ICML 2026 | | Subjects: | Artificial Intelligence (cs.AI) | | MSC classes: | 68T05 | | Cite as: |arXiv:2606.24842[cs.AI] | | | (orarXiv:2606.24842v1[cs.AI] for this version) | | |https://doi.org/10.48550/arXiv.2606.24842<br>Focus to learn more<br>arXiv-issued DOI via DataCite |
Submission history
From: Yikai Lu \[view email]
[v1] Tue, 23 Jun 2026 17:21:09 UTC (5,523 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled World Models in Pieces: Structural Certification for General Agents, by Yikai Lu and Yifei Wu and Xinyu Lu and Tongxin Li
[view license](http://arxiv.org/licenses/nonexclusive-distrib/1.0/ "Rights to this article")
Current browse context:
cs.AI
[< prev](https://arxiv.org/prevnext?id=2606.24842&function=prev&context=cs.AI "previous in cs.AI (accesskey p)") \| [next >](https://arxiv.org/prevnext?id=2606.24842&function=next&context=cs.AI "next in cs.AI (accesskey n)")
Change to browse by:
References & Citations
export BibTeX citation
Bookmark

Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer _(What is the Explorer?)_
Connected Papers Toggle
Connected Papers _(What is Connected Papers?)_
Litmaps Toggle
Litmaps _(What is Litmaps?)_
scite.ai Toggle
scite Smart Citations _(What are Smart Citations?)_
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv _(What is alphaXiv?)_
Links to Code Toggle
CatalyzeX Code Finder for Papers _(What is CatalyzeX?)_
DagsHub Toggle
DagsHub _(What is DagsHub?)_
GotitPub Toggle
Gotit.pub _(What is GotitPub?)_
Huggingface Toggle
Hugging Face _(What is Huggingface?)_
ScienceCast Toggle
ScienceCast _(What is ScienceCast?)_
Demos
Demos
Replicate Toggle
Replicate _(What is Replicate?)_
Spaces Toggle
Hugging Face Spaces _(What is Spaces?)_
Spaces Toggle
TXYZ.AI _(What is TXYZ.AI?)_
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower _(What are Influence Flowers?)_
Core recommender toggle
CORE Recommender _(What is CORE?)_
- Author
- Venue
- Institution
- Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.
Which authors of this paper are endorsers?\| Disable MathJax (What is MathJax?)
来源地区
Global
热度分
81
分类
研究进展
语言
en
