Trust Needs Evidence: Principles for Independent AI Evaluation

Independent evaluation is infrastructure for trust. Consequential claims should be backed by evidence that can withstand independent scrutiny.

Claims that an AI system is safe, accurate, fair, or better than its predecessors depend on measurement, and on choices about what to measure. AI measurement can take many forms: standardized tests, attempts to elicit failures, audits of system behavior, and monitoring how a system performs after deployment. The science behind it has made real progress, but it is young and unsettled, while AI systems are already being deployed and relied on. Like other fields where testing carries consequences, this one will mature by doing more of it: more evaluations, compared openly, using methods others can check and improve. That is a reason to build independent evaluation now, not to wait.

Done well, independent evaluation lets organizations deploy with confidence, gives developers a credible way to demonstrate quality, and catches problems when they are cheaper to fix. It is also a competitive advantage: countries and companies whose claims can be independently verified are better placed to deploy at scale and compete globally.

Credibility requires independence. In aviation, pharmaceuticals, and financial auditing, consequential claims do not rest solely on evidence controlled by the organization making the claim. Evaluators need enough autonomy in methods, access, technical judgment, funding, and reporting for their findings to be credible. Independence does not require isolation from developers, and independent evaluation complements rather than replaces developers’ own testing. Some AI developers have begun giving outside evaluators deeper access. That is welcome. The next step is to make such evaluation durable rather than discretionary. 

Independent evaluation does not require agreement on every AI policy question. It creates a common evidentiary foundation on which developers, users, regulators, and the public can make better decisions.

We call on governments, funders, developers, and standards bodies to create the conditions for independent AI evaluation that is rigorous, reliable, and credible to the public. Six conditions are especially important:

1. Fund the science, and the infrastructure it runs on, through structures that protect evaluator independence.

Invest in measurement science: what a test actually measures, how much of a reported difference is real rather than noise, and whether scores predict performance after deployment. Build a technically capable evaluation community, with methods that can withstand scrutiny. Invest in shared infrastructure: secure testing environments, standard test setups, and evaluations designed to resist contamination and gaming. Funding should be durable and structured so that developers cannot reward, punish, or defund evaluators based on their findings, or otherwise use funding to compromise their independence.

2. Make systems traceable, and failures reportable. 

AI systems used in consequential settings should be accompanied by enough information to investigate significant failures, including the ability to identify the system version, configuration, and conditions involved, subject to appropriate privacy and security protections. Post-deployment incidents are also evidence: they help reveal where evaluations missed important risks and whether test results predict real-world performance. Good-faith incident reporting should be supported by appropriate confidentiality and anti-retaliation protections while allowing findings and lessons to inform evaluation and oversight. 

3. Measure what matters to the people affected by AI systems, not just capability.

What we choose to measure shapes what we are able to see, compare, and improve. Evaluation should therefore examine not only what a system can do under test conditions, but whether safeguards work, how systems perform after deployment, and what effects they have in real-world use. Choices about what to measure should be made transparently and informed by people who use or are affected by these systems, domain experts, civil society, developers, and evaluators, sector by sector and in proportion to the stakes.

4. Provide the access necessary for credible evaluation, and make its terms transparent.

For sufficiently consequential systems, credible evaluation may require secure, proportionate access before deployment, not only to finished products after launch. Evaluators need enough access to test both what a system can do and whether its safeguards work. Access can be tiered to protect security, privacy, intellectual property, and other legitimate interests, without becoming a barrier to scrutiny itself. Restrictions that could affect the credibility of an evaluation, including limits on what evaluators can test or report, should be disclosed with the results so those results can be interpreted appropriately.

5. Protect responsible, good-faith testing. 

Evaluators¹ conducting responsible, proportionate, good-faith testing need clear protection against legal or contractual penalties arising solely from legitimate testing and responsible reporting. Existing protections and enforcement policies for good-faith security research provide a useful precedent. AI evaluation needs similarly clear safe harbors, while preserving safeguards against harmful or abusive conduct.

6. Make results comparable across borders.

Evaluation regimes should support international comparability through shared standards where appropriate, clear documentation of methods and uncertainty, and pathways for recognition where standards and safeguards are sufficiently compatible. Internationally recognized results can reduce duplicative testing, support coordination, and expand access to global markets.

¹An embedded evaluator may be protected by contract. The independent researcher who tests a deployed system from outside, without a badge or an agreement, is not. Contractual protection for the people inside does not substitute for legal protection of the people outside

Who’s Already Signed On

Signatories are participating in their personal capacities. Their participation does not represent or imply the involvement, endorsement, or position of any organization with which they are affiliated. Organizational affiliations are listed for identification purposes only.

Samir Jain, Center for Democracy & Technology

〰️

Vinh Nguyen, Council on Foreign Relations

〰️

Hoda Heidari, Carnegie Mellon University

〰️

Elham Tabassi, former Chief AI Advisor at NIST

〰️

Jason Pielemeier, Global Network Initiative

〰️

Tom Schaul, Google DeepMind

〰️

Suresh Venkatasubramanian, Brown University

〰️

Alondra Nelson, Institute for Advanced Study

〰️

Sarvesh Gupta, Oracle

〰️

Amy Chang, Cisco Systems

〰️

Stuart Russell, UC Berkeley

〰️

Amir Banifatemi, AI Commons

〰️

Tom Wheeler, FCC

〰️

Dr. Joy Buolamwini, Algorithmic Justice League

〰️

Laura Weidinger, Google Deepmind

〰️

Samir Jain, Center for Democracy & Technology 〰️ Vinh Nguyen, Council on Foreign Relations 〰️ Hoda Heidari, Carnegie Mellon University 〰️ Elham Tabassi, former Chief AI Advisor at NIST 〰️ Jason Pielemeier, Global Network Initiative 〰️ Tom Schaul, Google DeepMind 〰️ Suresh Venkatasubramanian, Brown University 〰️ Alondra Nelson, Institute for Advanced Study 〰️ Sarvesh Gupta, Oracle 〰️ Amy Chang, Cisco Systems 〰️ Stuart Russell, UC Berkeley 〰️ Amir Banifatemi, AI Commons 〰️ Tom Wheeler, FCC 〰️ Dr. Joy Buolamwini, Algorithmic Justice League 〰️ Laura Weidinger, Google Deepmind 〰️

Add your name

We invite everyone working to build, evaluate, deploy, govern, or study AI to add their name. We welcome collaboration across industry, academia, civil society, standards bodies, and public institutions on what these principles require in practice.

By signing, you agree to have your name and title displayed publicly as a supporter. We will not share your personal information with third parties without your consent.