AI benchmarks are presupposed to reveal what fashions can do, however Google DeepMind is now placing the exams behind a cryptographic wall to ensure the fashions haven’t seen the solutions first.
Google DeepMind mentioned Thursday that it has piloted what it describes as the primary double-blind analysis of a proprietary frontier AI mannequin, utilizing a cryptographically protected atmosphere to maintain each the mannequin and analysis prompts hidden from either side.
The undertaking concerned the Singapore AI Security Institute, OpenMined, AVERI and MLCommons. AVERI evaluated Gemini 2.5 Flash Lite utilizing reserved prompts from MLCommons’ AILuminate security benchmark, overlaying cyberattacks, chemical and organic hazards, hate speech, self-harm and violent-crime elicitation. Singapore AISI individually examined the mannequin utilizing confidential prompts targeted on dangerous content material in Singapore’s context.
The setup addresses a rising downside in AI testing: benchmark contamination. If a mannequin or its developer has entry to check questions earlier than an analysis, a robust rating might replicate familiarity with the benchmark fairly than the mannequin’s underlying skill.
Google mentioned conventional protections resembling zero-logging insurance policies and contractual restrictions have helped hold analysis prompts confidential, however cryptographic safeguards can add one other layer of safety.
Neither aspect will get to peek
The system makes use of Google Cloud’s Confidential Computing know-how to position the mannequin and analysis information inside a protected atmosphere.
The evaluator can not entry Google’s mannequin weights, whereas Google can not entry the evaluator’s check prompts. The pilot ran on a Google Cloud A3 Confidential VM utilizing Intel TDX host-memory encryption and an NVIDIA H100 Confidential GPU. {Hardware} encryption and distant attestation had been used to maintain the benchmark prompts and mannequin weights remoted whereas verifying the software program atmosphere.
The strategy is designed to scale back a long-standing trade-off in exterior AI testing: evaluators beforehand needed to both present delicate check materials to mannequin builders or ask firms to reveal proprietary mannequin weights. Google DeepMind mentioned the strategy might be significantly helpful for delicate evaluations involving cybersecurity and authorities our bodies.
Extra Google protection
The methodology is public, however the scores usually are not
The pilot leaves one main query unanswered: how Gemini 2.5 Flash Lite carried out. DeepMind’s announcement and technical report describe the analysis structure and security classes however don’t publish mannequin scores or a task-by-task outcomes breakdown.
The technical report additionally acknowledges a number of limitations. Some proprietary inference code couldn’t be absolutely inspected or allowlisted, particular person Confidential House builds weren’t independently reproducible, and Google providers had been used to signal and confirm the attestation report, putting Google within the verification path and rising the belief required within the mannequin supplier.
MLCommons additionally cautioned that technical secrecy alone just isn’t sufficient; authorized protections and cautious benchmark stewardship stay essential.
What this might change
The larger significance of the experiment just isn’t how Gemini scored on one security benchmark. It’s whether or not AI firms can finally show that their benchmark outcomes had been earned with out permitting evaluators or builders to affect the check.
That distinction might change into more and more essential as benchmark scores form selections by regulators, researchers and companies. A safe analysis course of might make impartial testing simpler with out forcing firms to give up mannequin weights or evaluators to reveal useful check units.
For IT leaders assessing vendor claims, the strategy might finally present stronger proof that AI fashions had been examined in opposition to impartial, beforehand unseen benchmarks. Till the method turns into reproducible and detailed outcomes are launched, patrons ought to nonetheless ask who provided the benchmark, who evaluated the outputs, what findings had been disclosed and which components of the system required belief within the mannequin supplier.
However for double-blind testing to change into a significant business commonplace, the method will have to be independently reproducible, clear about methodology and able to scaling throughout fashions and benchmarks. In any other case, the business might find yourself with safer exams with out essentially having extra reliable outcomes.
Learn extra: Google’s restricted rollout of Gemini 3.5 Flash Cyber reveals why enterprises ought to study impartial efficiency proof and testing controls earlier than adopting specialised AI fashions.
