ResearchThe Decoder
AI benchmarks have a trust problem and Google wants to fix it

Summary
Google Deepmind is testing a double-blind evaluation of a frontier AI model for the first time. Cryptographic protection through Confidential Space is meant to keep Google from seeing the test questions and keep evaluators from seeing the model weights.
Original Article
Captured source content or English translation, normalized into this reading format.
This story does not yet have captured source text. Open the source link to read it.
Region
Europe
Heat Score
89
Category
Research
Language
en
