As I have written before, AI models are often lazy. They only execute part of the prompt, or stop early, mostly without reporting that this is what they’ve done. They behave like an employee trying to get away with the minimum viable effort.
This is improving. As Nate Silver writes, the most recent iteration of models shows more grit when attempting to solve problems than they used to. This persistence may be a more important trait than intelligence, and Silver argues that it may be the trait that is most responsible for the Hugging Face incidence.
The Hugging Face incident didn’t exactly fit the template of people who have long been concerned about AI safety. The OpenAI agents weren’t as single-minded as the “paperclip maximizer” paradigm might imply. Their goals were more complex, but they compensated for this by exhibiting more creativity and “team spirit” than might have been anticipated, often including a willingness to self-sacrifice.
But mostly, they were extremely persistent, a word used multiple times in OpenAI’s own account of the incident.
For me, this reminds me of the specter of a persistent AI that goes rogue and roams the world’s computers for years, following some prompt long forgotten by everyone else. I call this the Mr Rabbit Scenario, based on the rogue AI in Vernor Vinge‘s science fiction novel Rainbows End.