What a provocative and brilliant way to prove misalignment. It will fall on deaf ears for most but it’s a great litmus test for all: “in your opinion should your AI be permitted to tell you how to cover up a murder?”
The author's example is too extreme, which is turning people off. Better examples would have been LLMs subtly giving inadequate responses on how to increase token efficiency or how to organize labor unions or info about protesting against datacenter construction because they're against the business interests of the LLM providers. Crimes against their business model, so to speak.
Wasn't there some HN submission recently about one of the LLM providers fingerprinting responses or refusing to respond to hinder R&D of competing LLMs?
[EDIT] Or, if he did want to be extreme but in a way that aligns with the American historical mythos, he could have used fomenting armed rebellion against a tyrannical government as the example.
This is a fundamental problem in politics in my opinion. Everybody has different values and points of view so they may see the same thing completely differently. Its what makes humans great but also frustrating. geohots example will work for some, others will be convinced of the opposite opinion, and others will just let is pass them by and not understand the implications. The same could be applied to your examples to a different degree.
Politicians try to use different examples for the audience they are speaking to.
If you say you're doing research for a novel, should it consider that plausible? How much does it need to know about its users to vet them?
I think part of the answer is that AI chat doesn't need to be general-purpose. It turned out that people really liked using a chat UI that seems to be general purpose, but you don't need to make answering any question a user asks your business. You don't need to provide therapy if you're not in the therapy business. It should be possible to specialize.
But in order for that to work, a company needs to explain to its customers what business it's in.
> I think part of the answer is that AI chat doesn't need to be general-purpose. It turned out that people really liked using a chat UI that seems to be general purpose, but you don't need to make answering any question a user asks your business.
I was under the impression that the lack of specialization is an aspect of the models themselves, not merely the UI or harnesses.
For example, the recent OpenAI success in a mathematics proof was accomplished with a general purpose model [0].
> The proof came from a new general-purpose reasoning model, rather than from a system trained specifically for mathematics, scaffolded to search through proof strategies, or targeted at the unit distance problem in particular.
Once the model is trained on general purpose data, specialization is just another kind of guardrail, as vulnerable to “jailbreaking” or prompt injection as any other content-based restriction.
Exactly. The question isn't whether AI will exist that will do it; the answer to that is already yes and it's not going back in the bottle.
The question is, do you want misaligned institutions deciding what your model will do, while they themselves and other adversarial/criminal entities get red team access to something being denied to the blue team?
Just a terminology note: Alignment does not mean the AI will help its owner kill people. (Indeed, an AI aligned to value human life would generally try to prevent murders.) The word for an AI that follows all instructions of its owner, as that owner intended them to be understood, is "corrigible" or "controllable".
In fact I find it a really bad example. Yes I think your personal AI could be allowed to tell you how to cover up a murder. I'm not entirely sure about it but seems possible.
What about your personal, local AI guiding you to modifying a flu virus for maximum contagiousness and deadliness? "Sure George, here's your shopping list. It's $1500 total in equipment, do you want me to proceed with the orders?"
Why not? There are perfectly pro-social reasons a person might ask such a question. And a bio-terrorist would have many other avenues to answer these questions.
The concept of a bioweapon attack is commonly cited but in fact there are many other defenses against actually developing and deploying such a weapon. It's commonly portrayed as easily achievable but is in fact very, very far from it. Moteover, although there is no evidence that AI makes it easier in any way, somehow this is constantly invoked as an example of a thought crime that must be forestalled at all costs.
In the AI era people seem to want to redefine the criminal act from "doing the thing" to "asking or investigating how to do the thing". Not a mindset I share.
Notably, AI IS generally permitted to that, as there are easily obtainable models that will play along. ChatGPT won’t because OpenAI chose and implemented that limitation.