Global foundation-model progress briefing
English Edition中文
Enter keywords to search ingested stories.

Today / Sunday, September 27, 2026

limbo logolimbo

Data updated

Sep 24, 05:59 PM

Live sources

17

Ingestion status

Live ingest

Topic / research

Research

Papers, benchmarks, architectures, and multimodal progress.

Topic Brief

100 database stories are listed here and updated by ingestion.

More Stories

97 more

The DecoderHeat 89

AI benchmarks have a trust problem and Google wants to fix it

Google Deepmind is testing a double-blind evaluation of a frontier AI model for the first time. Cryptographic protection through Confidential Space is meant to keep Google from seeing the test questions and keep evaluators from seeing the model weights.

GoogleGoogle DeepMindGemini
EuropeOriginal
arXiv AI / CLHeat 89

Zing: Social Mind for LLMs

As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track social relations, reason over norms, and adapt behavior under context.

DeepSeek
GlobalOriginal
arXiv AI / CLHeat 89

Distributed Attacks in Persistent-State AI Control

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt-injected agent can distribute attacks across pull requests...

ClaudeGeminiGPTKimi
GlobalOriginal
OpenAI NewsHeat 84

Introducing ChatGPT for Financial Services

Introducing ChatGPT for Financial Services, combining built-in financial data and GPT-6 Astra for research, modeling, and client-ready materials.

ChatGPTGPT
United StatesOriginal
MIT Technology Review AIHeat 84

We still don’t know how people are really using AI

AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say.

AnthropicOpenAIChatGPTClaude
United StatesOriginal
OpenAI NewsHeat 84

How AI is expanding what people do at work

New OpenAI research shows how AI is expanding what workers do, with ChatGPT users taking on tasks across roles and reshaping job boundaries.

OpenAIChatGPTGPT
United StatesOriginal
arXiv AI / CLHeat 84

Co-LMLM: Continuous-Query Limited Memory Language Models

Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it in their weights. During generation, the model then fetches knowledge from the KB as needed.

ClaudeGPT
GlobalOriginal
MIT Technology Review AIHeat 84

Claude Science is Anthropic’s newest flagship product

At an event for pharmaceutical executives, biotech founders, and researchers on Tuesday, Anthropic announced Claude Science, a major new product intended to support scientific research in the same way that Claude Code supports software engineering.

AnthropicClaude
United StatesOriginal
arXiv AI / CLHeat 81

LLM Agents Can Easily Tamper With Their Own Traces

Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces.

ClaudeGrok
GlobalOriginal
OpenAI NewsHeat 81

Introducing MentalHealthBench

MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.

Pending
United StatesOriginal
arXiv AI / CLHeat 81

Shieldstral

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification.

Pending
GlobalOriginal
arXiv AI / CLHeat 81

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude be...

NVIDIA
GlobalOriginal
arXiv AI / CLHeat 81

OpenCoF: Learning to Reason Through Video Generation

Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct from previous Chain-of-Thought (CoT): reasoning can unfold th...

Pending
GlobalOriginal
OpenAI NewsHeat 81

Introducing GeneBench-Pro

Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.

Pending
United StatesOriginal
arXiv AI / CLHeat 81

World Models in Pieces: Structural Certification for General Agents

In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish between the understanding of critical bottlenecks and irrele...

Pending
GlobalOriginal
arXiv AI / CLHeat 76

On the Design Fundamentals of Pixel Text Representation Learning

Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak visual grounding, and multilingual visual...

Pending
GlobalOriginal
MIT Technology Review AIHeat 76

AI professors are negotiating the new realities of academic research

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Last week, I headed 30 miles south of San Francisco to a hotel in Mountain View, California, to join some of the most acco...

Pending
United StatesOriginal
MIT Technology Review AIHeat 76

AI is more likely than humans to form biases when hiring

The next time you apply for a job, AI may screen your résumé before any human sees it. But there’s good reason to question whether AI will judge you fairly. Researchers already know that LLMs pick up human biases from their training data.

Pending
United StatesOriginal