全球大模型进展新闻浏览站
中文头版English
输入关键词,快速查找已抓取新闻。

今日版 / 2026年8月12日星期三

limbo logolimbo

数据更新时间

6月11日 11:00

启用来源

17

抓取状态

真实抓取

研究进展MIT Technology Review AI

谷歌DeepMind担忧数百万AI代理交互风险

摘要

谷歌DeepMind资助研究数百万不同AI代理在线交互的潜在危险。该公司AGI安全与对齐研究负责人罗欣·沙阿表示,大规模市场出现无需人类监督即可执行任务并遵循其他代理指令的代理,可能带来风险。

背景解释

随着AI代理技术发展,未来可能出现大量自主代理在网络上交互,这可能导致不可预测的行为或安全漏洞。DeepMind提前研究此类风险,有助于制定安全准则,确保AI系统可靠可控。

原文译文

以下为抓取到的原文内容译文,已统一为站内阅读格式。

阅读原文

Skip to Content

EXECUTIVE SUMMARY

Google DeepMind isfunding researchinto the potential dangers of situations where millions of differentAI agentsinteract with each other online.

According to Rohin Shah, who directs the company’s AGI safety and alignment research, the mass-market arrival of agents that can carry out tasks without human oversight and follow instructions given to them by other agents creates awhole new class of risk.

In an effort to address this, Google DeepMind—which made agent-based tools acenterpiece of Google I/O last month—has teamed up with several other organizations to announce a $10 million funding pot for researchers to study the behavior of multi-agent systems and come up with ways to prevent unsafe scenarios. Joining Google DeepMind are Schmidt Sciences, a philanthropic foundation set up by Eric and Wendy Schmidt; ARIA, theUK government’s moonshot agency; the Cooperative AI foundation, a UK-based nonprofit research outfit; and Google’s charitable arm, Google.org.

I asked Shah and James Fox, who leads the Science of Trustworthy AI program at Schmidt Sciences, what they hope to achieve with that $10 million. It’s no small sum, but it’s dwarfed by the budgets commanded by Google DeepMind’s own research teams.

The aim is to kick-start research outside tech companies, says Shah: “The strength of academia is that it can look really quite far into the future and do the kind of work that isn’t top of mind at industry labs.”

“The main issue is that there just isn’t really a field of research for multi-agent safety yet,” he adds. “And we would like there to be.”

The concern is that as more and more AI agents get deployed and begin working together, we could hit a tipping point where imagined scenarios become real. “We see this with humanity, too,” says Shah. “Our institutions can accomplish things that no individual human can.”

Shah thinks we have a few more months to go before agents are deployed throughout the economy in numbers that make potential risks a real concern. He wants to get ahead of that moment.

Risky business

What risks are we talking about, exactly? The possibilities that Shah and Fox have in mind mostly boil down to supercharged versions of bad things that happen on the internet already: scams, prompt injections (where an AI agent is fed malicious instructions, turning it into a self-guiding piece of malware), other forms of cyberattack. We look at what humans do now and ask what the agent version of that would be, says Shah.

“We’ve got this digital commons that is integral to how society works, and you really want to ensure that this doesn’t descend into just absolute anarchy,” says Fox.

(I asked Shah if they were considering any worst-case scenarios more on the doomer end of the spectrum, such as widespread economic collapse. “Certainly not if we’re talking by the end of the year,” he said. That’s only six months away! He laughed. “Okay, a while after that.”)

Shah and Fox both think that the only way to understand what might happen when large numbers of multi-agent systems interact with each other is to run realistic simulations. They want researchers to drop AI agents into sandboxes and study what they do.

You can’t predict what’s going to happen by studying single agents, or even small groups of agents, in isolation. You can’t assume that AI agents underpinned by LLMs will always act rationally, says Fox. And the complexity comes from having huge numbers of interactions at once.

Some researchers, including ateam at Google DeepMind, have argued thatartificial general intelligence(if possible at all) could come not from a single super-smart model but from a kind of agent hive mind, where the capabilities of the whole add up to more than the sum of its parts.

content frame

An error has occurred

![](https://www.technologyreview.com/subscribe/?itm_source=in-article&itm_medium=onsite&itm_campaign=InArticle-JulyAugust26-Issue_Current&utm_content=JA26ISSUE_REG_WEBUNITS&_ptid=%7Bkpdx%7DAAAAvne7HcPzgwoKV1VPQ05TVWdwdRIQbXIxd3FscGh6NTF3ZGVkeRoMRVhHRUk2ODdZME85IiUxODA1bmowMGRvLTAwMDAzN290NDdpYmM5NjZoZ24wMzI5MDlzKhpzaG93VGVtcGxhdGVLRlRaTVZCNFYyU1I0OTABOgxPVDFVS0dHWURBQkdCDU9UVk0wWjFFV0VGQlNSEnYthADwFm12OTcyODdkNloLNS4xODMuOTEuOTZiA2RsY2jc7pjSBnAteAQ)

New issue release + bonus AI content

Subscribe to access our latest issue and discover the breakthroughs shaping AI, climate, and modern engineering.

CLAIM OFFER

Lack of trust

Google DeepMind is not the only top AI firm warning about the risks of the technology it is building. A couple of weeks ago, Anthropic publishedguidelines for deploying AI agentsbased on an approach to cybersecurity known as zero trust, which starts with the assumption that a computer system is vulnerable, an agent is an attacker, and a breach will happen.

Refael Angel, cofounder and CTO of Akeyless, a cybersecurity firm based in Tel Aviv, agrees that understanding the new risks introduced by agent-based systems is crucial.

Every approach to security in the past has assumed that the machine in question was software written by a human, doing fixed things on fixed paths, says Angel: “An agent breaks all of those assumptions. It reasons, it improvises, and it can be hijacked by a single sentence buried in a document it was asked to read.”

Angel welcomes this new funding. “No single lab should author the safety standards everyone else has to trust,” he says. But he cautions that safety researchers can overlook boring problems that are already here in favor of more exotic hypothetical ones.

And yet, Fox notes, risks that were hypothetical a few years ago are now very real: “The future’s come more quickly than perhaps expected.”

hide

Deep Dive

Artificial intelligence

A startup claims it broke through a bottleneck that’s holding back LLMs

Subquadratic has now shared more details about its new model. But some are still skeptical.

By

Musk v. Altman week 1: Elon Musk says he was duped, warns AI could kill us all, and admits that xAI distills OpenAI’s models

Musk kept his cool, and OpenAI’s lawyer bulldozed him with piercing questions about his motivations for suing the company.

By

A reality check on the AI jobs hysteria

What do the numbers really say about the impact of artificial intelligence on the labor market? The answer might surprise you.

By

Anthropic’s Code with Claude showed off coding’s future—whether you like it or not

As tools like Claude Code get better, more and more developers are happy to hand off coding tasks to them. The way software gets built has changed for good.

By

Stay connected

Illustration by Rose Wong

Get the latest updates from MIT Technology Review

Discover special offers, top stories, upcoming events, and more.

Enter your email

Privacy Policy

Thank you for submitting your email!

Explore more newsletters

It looks like something went wrong.

We’re having trouble saving your preferences. Try refreshing this page and updating them one more time. If you continue to get this message, reach out to us at [customer-service@technologyreview.com](mailto:customer-service@technologyreview.com) with a list of newsletters you’d like to receive.

content frame

An error has occurred

Don't leave yet!

Save 25% + get bonus AI content

Subscribe to explore how engineering challenges are pushing the boundaries of human innovation in our new issue.

CLAIM 25% OFF

content frame

An error has occurred

This is a subscriber exclusive story.

Subscribe for full access to the July/August issue and lock in bonus AI content.

  • Monthly Digital

$12/month

CLAIM OFFER

  • 1 new digital issue bimonthly (6 yearly)
  • Unlimited access to our website and app
  • Exclusive access to the magazine archives
  • 20% discount on all events + access to our new event series, _Roundtables_
  • Best Offer

Digital + Print

$120/year

CLAIM OFFER

  • 6 new digital issues each year
  • Unlimited access to our website and app
  • Exclusive access to the magazine archives
  • 20% discount on all events + access to our new event series, _Roundtables_
  • Digital

$80/year

CLAIM OFFER

  • Everything you get with Digital, plus 6 new print issues each year

Sign in

来源地区

United States

热度分

73

分类

研究进展

语言

en