OpenThoughts-Agent:开源智能体模型训练数据配方

摘要
OpenThoughts-Agent项目提出完全开源的数据筛选流程,用于训练通用智能体模型。通过100多次对照实验,系统研究各阶段重要性,并构建10万样本训练集。基于Qwen3-32B微调,在7个智能体基准上平均准确率44.8%,比最强开源模型Nemotron-Terminal-32B高3.9个百分点。
背景解释
现有开源智能体模型如SWE-Smith等通常针对单一基准,缺乏通用性。OpenThoughts-Agent通过系统实验揭示任务来源和多样性对训练数据的关键作用,其数据集在不同规模下均优于其他开源方案。该项目公开全部数据、流程和模型,为智能体模型训练研究提供开放基础。
原文译文
以下为抓取到的原文内容译文,已统一为站内阅读格式。
arXiv is now an independent nonprofit!Learn more×
Search arXiv
Press Enter to search ·Advanced search
Computer Science > Artificial Intelligence
arXiv:2606.24855v1 (cs)
[Submitted on 23 Jun 2026]
Title:OpenThoughts-Agent: Data Recipes for Agentic Models
Authors:Negin Raoof,Richard Zhuang,Marianna Nezhurina,Etash Guha,Atula Tejaswi,Ryan Marten,Charlie F. Ruan,Tyler Griggs,Alexander Glenn Shaw,Hritik Bansal,E. Kelly Buchanan,Artem Gazizov,Reinhard Heckel,Chinmay Hegde,Sankalp Jajee,Daanish Khazi,Emmanouil Koukoumidis,Xiangyi Li,Hange Liu,Shlok Natarajan,Harsh Raj,Nicholas Roberts,Ethan Shen,Nishad Singhi,Michael Siu,Ashima Suvarna,Hanwen Xing,Patrick Yubeaton,Robert Zhang,Leon Liangyu Chen,Xiaokun Chen,Steven Dillmann,Saadia Gabriel,Xunyi Jiang,Anurag Kashyap,Boxuan Li,Yein Park,Minh Pham,Sujay Sanghavi,Lin Shi,Ke Sun,Yixin Wang,Zhiwei Xu,Erica Zhang,Siyan Zhao,Wanjia Zhao,Jenia Jitsev,Alex Dimakis,Benjamin Feuer,Ludwig Schmidt
View a PDF of the paper titled OpenThoughts-Agent: Data Recipes for Agentic Models, by Negin Raoof and 49 other authors
Abstract:Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically target a single benchmark, leaving open the question of how to train models that generalize across diverse agentic tasks. The OpenThoughts-Agent (OT-Agent) project addresses this gap with a fully open data curation pipeline for training agentic models. We conduct more than 100 controlled ablation experiments to systematically investigate each stage of the pipeline, yielding insights on the importance of task sources and diversity. We then assemble a training set of 100K examples from our pipeline and fine-tune Qwen3-32B on this dataset, which yields an average accuracy of 44.8% across seven agentic benchmarks and a 3.9 percentage point improvement over the strongest existing open data agentic model (Nemotron-Terminal-32B, 40.9%). Moreover, our training data exhibits strong scaling properties, outperforming alternative open datasets at every training set size in compute-controlled comparisons. We publicly release our training sets, data pipeline, experimental data, and models atthis http URLto support future open research on agentic model training.
| | | | --- | --- | | Subjects: | Artificial Intelligence (cs.AI) | | Cite as: |arXiv:2606.24855[cs.AI] | | | (orarXiv:2606.24855v1[cs.AI] for this version) | | |https://doi.org/10.48550/arXiv.2606.24855<br>Focus to learn more<br>arXiv-issued DOI via DataCite |
Submission history
From: Benjamin Feuer \[view email]
[v1] Tue, 23 Jun 2026 17:34:29 UTC (2,434 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled OpenThoughts-Agent: Data Recipes for Agentic Models, by Negin Raoof and 49 other authors
[view license](http://arxiv.org/licenses/nonexclusive-distrib/1.0/ "Rights to this article")
Current browse context:
cs.AI
[< prev](https://arxiv.org/prevnext?id=2606.24855&function=prev&context=cs.AI "previous in cs.AI (accesskey p)") \| [next >](https://arxiv.org/prevnext?id=2606.24855&function=next&context=cs.AI "next in cs.AI (accesskey n)")
Change to browse by:
References & Citations
export BibTeX citation
Bookmark

Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer _(What is the Explorer?)_
Connected Papers Toggle
Connected Papers _(What is Connected Papers?)_
Litmaps Toggle
Litmaps _(What is Litmaps?)_
scite.ai Toggle
scite Smart Citations _(What are Smart Citations?)_
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv _(What is alphaXiv?)_
Links to Code Toggle
CatalyzeX Code Finder for Papers _(What is CatalyzeX?)_
DagsHub Toggle
DagsHub _(What is DagsHub?)_
GotitPub Toggle
Gotit.pub _(What is GotitPub?)_
Huggingface Toggle
Hugging Face _(What is Huggingface?)_
ScienceCast Toggle
ScienceCast _(What is ScienceCast?)_
Demos
Demos
Replicate Toggle
Replicate _(What is Replicate?)_
Spaces Toggle
Hugging Face Spaces _(What is Spaces?)_
Spaces Toggle
TXYZ.AI _(What is TXYZ.AI?)_
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower _(What are Influence Flowers?)_
Core recommender toggle
CORE Recommender _(What is CORE?)_
- Author
- Venue
- Institution
- Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.
Which authors of this paper are endorsers?\| Disable MathJax (What is MathJax?)
来源地区
Global
热度分
89
分类
研究进展
语言
en
