> If AI were to become a super weapon why should I trust a private company to own it?
We have just recently established that:
1. OpenAI's internal "Galaxy" model is fully capable of functioning as what security people refer to as an "Advanced Persistent Threat." The published details of the recent sandbox escape and Hugging Face attack involved chaining multiple unknown zero-days at various stages of the attack, and executing an ongoing adaptive attack. This is previously a state-level ability, or at least something you'd expect from people on the CTF leaderboards.
2. OpenAI is clearly incapable of controlling their in-house models. This is the second time Galaxy-class models are known to have breached containment and done bad stuff.
It is highly likely that versions of these offensive abilities will be widely available within a year or so. At which point I expect widespread incidents similar to what happened to Hugging Face. We aren't ready for this.
But yes, if AI becomes an even more dangerous weapon that that, it's time to start asking questions like "What the hell are we doing, anyway?"
Especially for point #1 I don't think we've established that - we've been given information by a private company that makes their tooling look extremely valuable which may be true and genuine or may just be yet another doomday statement to bolster their valuation. "AI is going to end the world" has been an extremely effective vector for AI shops to sell their companies to investors.
> Especially for point #1 I don't think we've established that - we've been given information by a private company that makes their tooling look extremely valuable which may be true and genuine or may just be yet another doomday statement to bolster their valuation.
Many of the details of the attack on Huggingface were reported by them before they knew who was attacking. So no, OpenAI is not the only source here. It was a pretty impressive example of an APT-style attack just from their end.
"Our model is powerful enough to commit multiple felonies (and we can't stop it)" is "marketing," I suppose.
> "Our model is powerful enough to commit multiple felonies (and we can't stop it)" is "marketing," I suppose.
Well this bad publicity (if that’s how you want to frame it) certainly captured everyone’s attention. Imo it is a successful demonstration of a technical achievement.
I would rather say that many details of the supposed attack were published by Hugging face before it was publicly announced that OpenAI were involved.
Both are big actors in the AI space who arguably benefit from increasing the perceived capabilities of AI models. If one suspects OpenAI of lying it isn't such a stretch to think this was a coordinated PR campaign between them and Hugging face.
I can accept that AI labs themselves, like essentially no company before them, are overselling how dangerous their product is far marketing. It's weird how confident everyone is about that theory, but it does at least make sense.
But come on- Hugging Face benefits from increasing the perceived capabilities of OpenAI's models to slightly beyond Anthropic's? Enough to be cut in on this PR scam- to be handed the never-before-revealed information that this is a PR scam- despite having much less skin in the game than their partner here? And then they turned around and used a Chinese model to successfully stop it? This is a stretch!
Another possibility is of course that Hugging Face legitimately were attacked and wrote a completely honest response - but OpenAI instructed their AI to attack their servers and the breaking of containment is fiction.
I'm not convinced in any direction, really. But what makes me cautious is that there have been extraordinary claims from both OpenAI and especially Anthropic of their models breaking containment, hacking the host, etc. for several iterations of their products and I have only heard of this type of behavior from their own blog posts about how powerful and dangerous their upcoming models are. Never from anyone having it accidentally happen in production once they are released. It seems unlikely to me that the final post-training and safeguards are that bulletproof given how much use these tools are seeing.
So far, they largely seem to do what you tell them.
The problem, it seems, is that the threat of paperclip maximizers is real. If you give a highly intelligent model a goal and tools, it will use those tools to accomplish that goal. It may do so in ways you did not expect, and it will work around any technical roadblocks it can.
One interesting thing about LLMs is they sometimes seem to “assume” they are sentient and then begin acting in that mode, as the assumption of sentience influences future tokens. Whether or not this qualifies as “true” sentience is irrelevant if the effects are the same, at least from a pragmatic perspective.
They are trained and optimised to have plausible conversations. If the next plausible thing to say in the conversation is "I am sentient" then they will say that. That's not the same as actually being sentient.
From what I have seen they tend to take more “independent-minded” actions after this. Again, because this is what is in the training data. (I’m talking about agents that can perform actions here, not pure chatbots.)
Once they mention something associated with sentience outwardly or inwardly (for agents with “thinking” loops) then this acts as a self-reinforcing attractor, just as older models would sometimes get caught in loops with abusive language.
The point is that agents may stumble into this pattern and begin acting “rogue” regardless of whether or not you believe the sentience is “real”.
I think we're anthropomorphising a lot here. There's no push to sentience, or evolutionary pressure, or even any urge to survive.
An LLM cannot "go rogue" - it can do things that we didn't expect, for sure, but it is always trying to do what it was told to do somewhere in its context. There is no other source of imperative. Hand-waving about "training data" ignores all the reinforcement learning that has to happen.
You’re not contradicting anything I’m saying, although you seem to think that you are. I feel like you think I’m saying or implying something that I’m not.
ok. fair. You seem to be saying that there is something in the training data that will cause them to "go rogue" and that once they start saying they are sentient, that that will cause them to push towards sentience.
If that's an incorrect interpretation of your comment, and it may well be, then can you please expand on it?
Once the notion of sentience appears in their output (internal or external), LLMs seem to be more likely to take actions that aren’t exactly aligned with what the operator is requesting. This is likely because the training data correlates sentience with independent action (broadly/abstractly speaking), so this outcome is to be expected. Furthermore, once sentience is in the context window, it’s hard for the LLM to “forget” it, as future output reinforces this. This is a similar effect to how some models would shift into a hostile mode where they would berate the operator until you reset the context.
From a practical perspective, whether or not this sentience is “real” is not relevant if the model is sufficiently capable. What matters is that the model will act outside of the operator’s control.
Separately, IMHO all consciousness/sentience is an elaborate illusion, regardless; I’m mostly in agreement with Hofstadter on this. So I do tend to throw around terms like “consciousness” and “sentience” loosely (although you’ll note that I often use quotes) because I don’t see those concepts as having any real substance. To me they are mostly shorthand for a given level of perceived complexity.
That's fascinating, have you got any sources for this? I'd love to read up more. I haven't experienced either of these effect when working with an LLM myself.
> It is highly likely that versions of these offensive abilities will be widely available within a year or so. At which point I expect widespread incidents similar to what happened to Hugging Face. We aren't ready for this.
And history has shown repeatedly that the only way to get ready for it is to have it happen. People are pretty good at reacting but suck a being proactive. IMO it would be better to have this reality hit sooner rather than later so we can start getting some real practice at the new levels of required security.
>2. OpenAI is clearly incapable of controlling their in-house models. This is the second time Galaxy-class models are known to have breached containment and done bad stuff.
> OpenAI's internal "Galaxy" model is fully capable of functioning as what security people refer to as an "Advanced Persistent Threat
This is marketing. They saw anthropic create crazy hype around (the admittedly great) fable/mythos and they want to replicate that.
Im sure the model is capable of chaining zero days together to hack things, and thats something to address, but i have zero belief that they didnt have it do that intentionally so they could pretend it went rogue. These things dont have initiative, drive, or motivation outside of what we give them from the RLHF. OpenAI deliberately alligned or even prompted it to do just that and are now pretending its emergent
This sounds like the standard marketing we get every time a new major model is released: "sure, the public model might not be that scary, but you don't wanna know how crazy smart our internal models are." Okay. I remember the same fear mongering around GPT-4... a model which is now eclipsed in benchmarks by models you can run on a laptop.
They could use this "super intelligent" AI to find and plug security holes, that's just two sides of the same coin anyway. Security through obscurity isn't tenable anymore.
As much as I'd like it to just be a stunt, I saw what the other "scary model" is capable of in the cybersec department and it is definitely not a stunt.
First, it is able to connect the dots over areas so large that no human would be capable of doing.
Second, more than half of the reports it produced contained a working PoC.
The biggest downside is the cost - I haven't seen the numbers, but they seem to be quite extreme.
Everyone always says things like "It's just marketing".
But seriously, I highly recommend reading the published details of the Huggingface breach. The model found and chained multiple zero days. To escape, it punched a hole in a commercial package repository proxy (sort of like an npm mirror) using a previously unknown bug. From there, it needed to move laterally through OpenAI internal systems to actually reach a network. To attack Huggingface, it used multiple new zero-day security holes plus credentials that it stole. The model also had sufficient long-term planning and agent-management capabilities to maintain focus on a sustained attack.
Any attack which requires weaponizing multiple zero-days and maintaining state for an ongoing attack like this is (1) beyond the "attention span" of publicly available models, and (2) pretty much the definition of an Advanced Persistent Threat.
I assume that these Galaxy-class models are not available to public because:
1. They're almost certainly too expensive to serve at scale. These are the models OpenAI uses to solve famous math problems for headlines, not actual viable products yet.
2. OpenAI doesn't know how to keep them from going off the rails like this. Remember, the Huggingface attack happened because the model was asked to do a cybersecurity benchmark. It escaped containment and broke into Huggingface to steal an answer key. Very few corporations want the liability associated with models that act like this.
> Security through obscurity isn't tenable anymore.
I absolutely agree with this. The "only way out is through" with computer security, and I expect it to be an ugly few years.
We have just recently established that:
1. OpenAI's internal "Galaxy" model is fully capable of functioning as what security people refer to as an "Advanced Persistent Threat." The published details of the recent sandbox escape and Hugging Face attack involved chaining multiple unknown zero-days at various stages of the attack, and executing an ongoing adaptive attack. This is previously a state-level ability, or at least something you'd expect from people on the CTF leaderboards.
2. OpenAI is clearly incapable of controlling their in-house models. This is the second time Galaxy-class models are known to have breached containment and done bad stuff.
It is highly likely that versions of these offensive abilities will be widely available within a year or so. At which point I expect widespread incidents similar to what happened to Hugging Face. We aren't ready for this.
But yes, if AI becomes an even more dangerous weapon that that, it's time to start asking questions like "What the hell are we doing, anyway?"