
It added: “Third party assessments are most useful when they address specific, consequential questions: Does the evidence support a lab’s safety case and safety claims? Do evaluations adequately test the risks they are intended to measure? Do safeguards work under realistic conditions?”
Nothing about enforcement
Pieter Arntz, a malware intelligence researcher at Malwarebytes, said that giving third parties rules for engagement is certainly a good thing, but the implication behind such rules is that they are somehow enforceable. And the document says nothing about that enforcement process.
“It sets some useful expectations for independence and rigor, but it does not itself compel OpenAI to submit to a particular scope, publish adverse findings, or change deployment decisions,” he pointed out. “Its credibility will depend on the terms of individual assessments, and what outsiders are allowed to see. Its value therefore hinges on whether OpenAI accepts genuinely inconvenient scrutiny, and whether results, redactions, remediation, and deployment decisions can be independently checked.”