Yes. Both the protocols and the technical means are here to prevent unintended usage and to enforce the intended one. It is the same thing. The ring (and the mechanism behind it) on a hand grenade is both the safeguard and the trigger.
But apart from the terminological nitpicking. The safeguards on the Soviet nuclear arms were installed by the Soviets under the military pressure from the US.
How are you going to install safeguards on Chinese, Russian, … AIs?
I don’t think that is true. The pressure was on the military to make sure they only released the missiles when their side could win. Not when the resulting havoc had no possible winner. Not when releasing the missiles was going to kill their side too.
That is again relevant for AI: control it, because letting go will hit your side too.
Just as a reminder: The USSR signed the Biological Weapons Convention (1972), which banned development, production and stockpiling of bioweapons. And continued their development and production all the way.
I told my students they were not allowed to use AI in my course, but that it went both ways: I also would not use AI at all. It’s the closest I’ve been to a standing ovation in all my years teaching (computer/electrical engineering).
There are a few of us. I always try to remember that repeating a bad argument a million times does not make it good.
Personally, I think climate change might have direr consequences for humanity in a shorter time frame than AI. And all those datacenters don’t help, which is a good reason not to rely on AI so much.
“… in other areas, such as tests to see if it will take destructive actions when pushed, Astra hides much of its thinking (see chart). And it does this most when it is made aware of being monitored, according to Tomek Korbak, who works on safety at OpenAI.”
Key phrase is “according to Anthropic.” I’m sure there are biologists who think Claude shouldn’t be used for gain-of-function research, but it is not inherently malicious even in military labs. It’s fair to deter VPN users from legally blocked countries, but it’s really looking like Anthropic exaggerated that to “potential bioweapons” as a public show of their self-regulation and does not generally discourage the use of Claude for gain-of-function research. In reality, gain-of-function research has led to treatments, not bioweapons attacks, and AI software suites alone won’t change that.
What would change that is a distraction from bringing regulations of the existing harms in the AI age (as opposed to the Terminator scenario) up to par with regulations of other potentially dangerous technology. For example, if AI-generated summaries with the extra risk of hallucinations were allowed as a substitute for human-written summaries (with AI research assistance or not) in critical situations:
I’d add that the so-called peace movement in the West was definitely sponsored by the USSR, and the (hugely “successfull” here in Germany) anti-nuclear movement was most probably sponsored by Russia, too. Though I wouldn’t go as far as to maintain we caused Chernobyl just for that purpose.
Google disclosed that because they could claim Gemini was controlled. All cases involved extrapolating from information leaked through public repos, not discovering zero-days, and the model stopped on its own. “Hey everyone, our model is AGI too because it does things we don’t want, but it cares about the law!”
I feel like this is a bunch of hype for the new market of cyber security agents to roll out to protect everyone from the other flagship models. Use a abliterated model from qwen and you wont think the guardrails are a joke on the larger models
This is what I mean, flagship/frontier whatever you wanna call it, right now its a marketing campaign by each major ai company like gemini, openai and claude all had recent “hacking attempts” all aired on the news… this creates a need for a defense to that or no. Were seeing the prep stages for the next wave of agents to be defense agenst in my opinion, agents that watch processes, have root aceess and kill pid or monitor network traffic, like the Jev model does with almost instant speed