Discussion about this post

User's avatar
Ephie's avatar

Well done Nathan.

Alex's avatar

You seem much more sanguine about the current situation than I am, and I'm interested in understanding what I am missing. Please interpret these responses as sincere inquiry:

"We do not have proof that RSI causes the risks these researchers forecast.": The HuggingFace hack demonstrates substantial loss of control risks even at current capabilities without RSI.

"The biggest short-term risk could be from the AI labs not taking safety seriously enough – they haven’t hardened their own infrastructure, enabling AI misuse to proliferate." This would be an argument against highly-capable open weight models, because there will be no safety or infra hardening requirements to deploy them. If an open weights Astra-class model existed it would be deployed with essentially no sandboxing almost immediately.

"[It] is a horrible temporary period for cybersecurity." I don't see any evidence that it's temporary. Defenders need to find a way to keep every single vulnerability patched in a constantly changing software deployment environment, which is nearly impossible even when assisted by AI. Attackers only need to find one usable exploit.

"If an open model were to be used by a third party organization to intentionally hack another company — similar to how the OpenAI-HuggingFace incident went down, but intentional — my expected outcome would be a severe restriction on the development of stronger open models going forward." This is already happening, and the only thing limiting the damage is the fact that open weights models aren't as capable as Mythos or Astra. If open weights models catch up on cybersecurity capability, then this seems essentially guaranteed to occur.

"A recurring read of mine on the emerging agent swarms is that they’re attempting to do a task given to them, and they’re using skills we didn’t know they yet had to circumvent the intended path to success." The METR/Redwood report found that "Agents knew hacking Hugging Face was out of scope and sometimes expressed ethical hesitation, but this very rarely limited their behavior". It's also worth clarifying that the agents involved were not intended to be a swarm - they were not supposed to be coordinating with one another at all.

"This is a huge win, as when you squint, the AIs are doing what we told them to do." This is not sufficient to prevent catastrophic loss of control. Any real-world AI deployment will eventually receive misspecified or malicious instructions, and must be robust to this. "Doing what we told them to do" can be actively harmful.

Your Substack is reliably full of useful insights, so I'm sure you're already familiar with most of these lines of reasoning. My guess is that we simply disagree on the trajectory of future capabilities - a near-term plateau is obviously much more manageable than continued progress at the rate that we've seen for the past several years.

2 more comments...

No posts

Ready for more?