Global foundation-model progress briefing
English Edition中文
Enter keywords to search ingested stories.

Today / Wednesday, August 12, 2026

limbo logolimbo

Data updated

Aug 11, 05:04 PM

Live sources

17

Ingestion status

Live ingest

Topic / research

Research

Papers, benchmarks, architectures, and multimodal progress.

Topic Brief

100 database stories are listed here and updated by ingestion.

arXiv AI / CLHeat 89

Zing: Social Mind for LLMs

As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track social relations, reason over norms, and adapt behavior under context.

DeepSeek
GlobalOriginal

More Stories

97 more

arXiv AI / CLHeat 89

Distributed Attacks in Persistent-State AI Control

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt-injected agent can distribute attacks across pull requests...

GPTClaudeGeminiKimi
GlobalOriginal
OpenAI NewsHeat 84

How AI is expanding what people do at work

New OpenAI research shows how AI is expanding what workers do, with ChatGPT users taking on tasks across roles and reshaping job boundaries.

OpenAIGPTChatGPT
United StatesOriginal
arXiv AI / CLHeat 84

Co-LMLM: Continuous-Query Limited Memory Language Models

Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it in their weights. During generation, the model then fetches knowledge from the KB as needed.

GPTClaude
GlobalOriginal
MIT Technology Review AIHeat 84

Claude Science is Anthropic’s newest flagship product

At an event for pharmaceutical executives, biotech founders, and researchers on Tuesday, Anthropic announced Claude Science, a major new product intended to support scientific research in the same way that Claude Code supports software engineering.

AnthropicClaude
United StatesOriginal
arXiv AI / CLHeat 81

Shieldstral

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification.

Pending
GlobalOriginal
arXiv AI / CLHeat 81

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude be...

NVIDIA
GlobalOriginal
arXiv AI / CLHeat 81

OpenCoF: Learning to Reason Through Video Generation

Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct from previous Chain-of-Thought (CoT): reasoning can unfold th...

Pending
GlobalOriginal
OpenAI NewsHeat 81

Introducing GeneBench-Pro

Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.

Pending
United StatesOriginal
arXiv AI / CLHeat 81

World Models in Pieces: Structural Certification for General Agents

In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish between the understanding of critical bottlenecks and irrele...

Pending
GlobalOriginal
MIT Technology Review AIHeat 76

AI professors are negotiating the new realities of academic research

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Last week, I headed 30 miles south of San Francisco to a hotel in Mountain View, California, to join some of the most acco...

Pending
United StatesOriginal
MIT Technology Review AIHeat 76

AI is more likely than humans to form biases when hiring

The next time you apply for a job, AI may screen your résumé before any human sees it. But there’s good reason to question whether AI will judge you fairly. Researchers already know that LLMs pick up human biases from their training data.

Pending
United StatesOriginal
arXiv AI / CLHeat 76

Real-time fall detection based on vision for low-power edge platforms

Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches predominantly frame it as static pose classification or discrete temporal pattern matching, fundamentally overlooking the instability dynamics...

Pending
GlobalOriginal
MIT Technology Review AIHeat 76

What Anthropic’s latest AI discovery does—and doesn’t—show

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Anthropic—currently the world’s most valuable AI company, with a nearly $1 trillion valuation—has a reputation for publish...

Anthropic
United StatesOriginal
MIT Technology Review AIHeat 76

Anthropic found a hidden space where Claude puzzles over concepts

The AI firm Anthropic has developed a technique that has given it the clearest glimpse yet at what’s really going on inside large language models as they answer questions or carry out tasks. What they found ranges from the mundane to the unnerving.

AnthropicClaude
United StatesOriginal
arXiv AI / CLHeat 73

Learning to Trace Seiberg Dualities

Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical systems. In practice, though, it can often be computationally challenging to establish when two systems are dual, even when all of the "rules o...

Google
GlobalOriginal
MIT Technology Review AIHeat 73

A fundamental flaw leaves LLMs strikingly vulnerable to attack

It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month.

Pending
United StatesOriginal
The DecoderHeat 73

Microsoft's open-weight AI push is so obviously an Azure play it hurts

Microsoft, along with Meta, Nvidia, and more than 20 other companies, is pushing for open-weight AI models in an open letter. The strategic logic is simple: the more models running on Azure, the less Microsoft depends on expensive OpenAI and Anthropic models.

OpenAIAnthropicMetaNVIDIA
EuropeOriginal
arXiv AI / CLHeat 73

FARS: A Fully Automated Research System Deployed at Scale

Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks.

Pending
GlobalOriginal