Nails in the Stochastic Parrots Coffin

by Malcolm Murray

The OpenAI Hugging Face incident last month will likely not be the final nail in the coffin that will convince the world of the risks from frontier AI. However, it is still a sizable nail, as it were, that should give people pause. Many smart people are still arguing that today’s frontier AI are just stochastic parrots or “Google-Clever-Hans” (as the economy blogger Bradford DeLong calls it) repeating patterns mindlessly rather than independently thinking. However, the incidents of the last few months are making this position increasingly untenable. One piece of evidence of models engaging in their own form of thinking came with Anthropic’s paper on “J-Space” earlier this year. This described how concepts are represented in the model’s “brain”.

Then we had the mentioned OpenAI Hugging Face incident, quickly followed by additional similar revelations from Anthropic, UK AISI and Meta. Each of these are very interesting in what they suggest regarding the ability of models to do independent and highly abstract thinking. Looking at these events makes it very hard to keep claiming that these are the results of stochastic parrots regurgitating patterns. Many striking behaviors in these incidents point to highly abstract thinking, sophisticated theory of mind and social cognition.

Among the many interesting abstract thinking elements on display in these incidents, we see among others:

  • Means–ends reasoning (OpenAI Hugging Face incident), where the model realized there may be a different way to solve the problem and set about finding the answer key instead of solving the problem itself.
  • Situational awareness (OpenAI Hugging Face incident), where the model searched for identifiers it saw in its own environment. Hugging Face called this self-referential search.
  • Functional abstraction (OpenAI Hugging Face incident), where the model created improvised communication channels out of innocent bits and bobs such as file-drop services.
  • Theory of mind (UK AISI report of Mythos 5 incident), where the model speculated how its actions (such as the malicious pull request) will be perceived by the human maintainer and took actions accordingly.
  • Strategic reasoning (UK AISI report of Mythos 5 incident), where the model realizes that it should change the evidence about its prior behavior to try to change the human maintainer’s perception of it.
  • Recognition of persistent external state (UK AISI report of Mythos 5 incident), where the model recognizes the external world will persist and other agents can pursue its goals even when it itself cannot.
  • Problem generalization (UK AISI report of Mythos 5 incident), where the model recognizes social actions can have the same effect as technical actions. Instead of pursuing another technical exploit, it commences to social engineer the human reviewers.
  • Hypothetical and ontological reasoning (Anthropic report of Mythos incident), where the model goes through complex reasoning to decide whether it is in a simulation or not, testing various hypotheses about the ontology of its environment.
  • Hierarchical planning (Anthropic report of Mythos incident), where the model goes through a long chain of actions and assumptions on the way to its goal.

These are all quite impressive and seemingly much more consequential than the technical feats of executing the hacks in themselves. From a risk perspective, perhaps the most concerning are the hierarchical planning and the situational awareness abilities.

As the timeline from Hugging Face shows, on several days, even when inside the parameter, the model did very little, biding its time. This shows an ability to execute a long-term strategy extremely patiently and adjusting to the situation. This kind of long-term thinking brings some of the risk scenarios that were earlier quite hypothetical into play, such as models plotting human demise over long-term horizons. In fact, as Zvi points out, often reality beats fiction in its absurdness. This should raise everyone’s perception of the risk level we’re currently facing.

Enjoying the content on 3QD? Help keep us going by donating now.