---
format: "aidr-story-markdown/v1"
id: "4b5cec038efd24b6df15ce8f642b1066027c070ecda95d3a744726fcc8808733"
canonical_url: "https://aidr.today/4b5cec03?lang=en"
title: "Google DeepMind Runs First Double-Blind AI Eval, Tackling Benchmark Contamination"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-27T14:57:37.000Z"
category: "Research"
topics: ["google","gemini","benchmark","safety","open-source"]
source_urls: ["https://huggingnews.com/ai/google-deepmind-runs-first-double-blind-ai-eval-tackling-benchmark-conta-a0e15d3a","https://x.com/GoogleDeepMind/status/2092961763553677387","https://x.com/Miles_Brundage/status/2092970514168140100"]
summary: "Google DeepMind and AVERI tested Gemini 2.5 Flash-Lite using a secure enclave to keep both model weights and test prompts private. The evaluation utilized the AILuminate safety benchmark from MLCommons in a double-blind process where neither the developer nor the evaluator had access to the other's sensitive assets. The project, which also involved OpenMined and MLCommons, seeks to eliminate benchmark contamination, a problem where AI models are trained on the very data used to test them. This production-scale test follows a 2024 pilot involving GPT-2 and the AI Security Institute, and is part of an AVERI strategy to establish open source auditing standards for frontier models."
---

# Google DeepMind Runs First Double\-Blind AI Eval, Tackling Benchmark Contamination

> [Open the canonical story](<https://aidr.today/4b5cec03?lang=en>)

**Published:** 2026-08-27T14:57:37.000Z
**Category:** Research
**Topics:** google, gemini, benchmark, safety, open\-source

## Summary

Google DeepMind and AVERI tested Gemini 2\.5 Flash\-Lite using a secure enclave to keep both model weights and test prompts private\. The evaluation utilized the AILuminate safety benchmark from MLCommons in a double\-blind process where neither the developer nor the evaluator had access to the other's sensitive assets\. The project, which also involved OpenMined and MLCommons, seeks to eliminate benchmark contamination, a problem where AI models are trained on the very data used to test them\. This production\-scale test follows a 2024 pilot involving GPT\-2 and the AI Security Institute, and is part of an AVERI strategy to establish open source auditing standards for frontier models\.

## Sources

- [Story source](<https://huggingnews.com/ai/google-deepmind-runs-first-double-blind-ai-eval-tackling-benchmark-conta-a0e15d3a>)
- [Story source](<https://x.com/GoogleDeepMind/status/2092961763553677387>)
- [Supporting source](<https://x.com/Miles_Brundage/status/2092970514168140100>)

