Embedded Evaluators: A Necessary but Perilous Step for AI Safety The recent proposal from Anthropic and OpenAI to embed third party evaluators within their systems has sent shockwaves through the tech industry.
At first glance, this development seems like a straightforward win win: AI companies get an independent check on their work, while researchers gain unprecedented access to the models they've been trying to understand.
However, scratch beneath the surface and you'll find a complex web of power dynamics, conflicting interests, and uncertain outcomes. Control is the elephant in the room.