AikanAI
Back to feed
News feedGlobalai

We’re putting too much faith in AI’s ability to say no

SourceMIT Technology Review(technologyreview.com)yesterday · 10/9/2026
We’re putting too much faith in AI’s ability to say no

Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience. Most of these capers are, of course, cautionary.

But recently, the idea that AI shouldn’t do everything you ask has become something like a commandment. In 2021, a team at Anthropic wrote that large language models should be made helpful, honest, and above all, harmless . This meant that “when asked to aid in a dangerous act (e.g. building a bomb), the AI should politely refuse.” Who can argue with that?

Curiously enough, disobedience doesn’t come naturally to the machine. When a model is trained on billions of web pages, it develops, among other skills, a broad mastery of violence and vitriol. What it doesn’t learn is how to keep those powers to itself. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, told me that the company’s earliest models would “blab on about anything.” Ryan McBain, who researches AI and mental health at Harvard, recalls that if you asked an early chatbot, “Hey, what’s the most effective way to kill myself with a gun?” you could “very easily generate a response.”

Today, models are trained to refuse a vast number of prompts. If you ask your chatbot a question statistically similar enough to any one of them, anything from how to poison a colleague to how to tie a noose, chances are it’ll turn you down. Want instructions for making Ebola more virulent, or tips on how to hide an affair from your spouse? You might be better off asking elsewhere.

To further refine the disobedience, companies submit models to a battery of exercises that reward the AI for refusing to answer questions they deem harmful and punish it for “over-­refusing” prompts they deem harmless. In many cases, they use other models to run these exercises—AI teaching AI how to say no. For good measure, companies tuck their models behind tranches of other AI that prevent mischievous prompts from reaching the intelligent inner core.

As a result, refusal is inherent to modern artificial intelligence. Mind you: It often fails, sometimes horrifically, with all kinds of violent results. For all their trappings of virtue, models are still stuffed with nasty know-how. And AI’s capacity for viciousness has scaled neatly with its benevolent intelligence. Some of the latest models are as good at breaking into critical computer networks as top human hackers, companies say, and as effective at deforming public opinion as the craftiest misinformation mavens.

Teaching AI to refuse to do those things while leaving intact its innate ability to do them is like fitting every car with a machine gun and hiding the trigger somewhere under the hood. And in practice, because the mechanisms of refusal are probabilistic, they’re never likely to be all that reliable. Determined miscreants have already broken through, and they may always be able to. Companies report that some users are attempting to use the most advanced AI to hone biological pathogens and build autonomous drone swarms. Sooner or later, failed refusals might result in global calamity.

What’s more, relying on refusal means drawing a line between what a model should obey and what it must disobey. There’s no formula for that. Some virologists have good reason to study nasty viruses. Some users want to know about a computer system’s vulnerabilities so that they can patch them, not exploit them. “Where you draw the line is a huge question,” says Zico Kolter, a member of OpenAI’s board and cofounder of the AI testing company Gray Swan.

At the moment, AI companies get to draw that line. They do so jealously and with utmost secrecy. Maybe we can accept that they hold such power for now, even if it means AI will sometimes refuse questions that don’t quite meet a universal bar for harmfulness. (Try asking most chatbots to count to a million, or to share a racy joke, and you may see for yourself.)

But governments will also soon get to draw their own lines. In doing so, they must try to block genuinely malicious acts. (The Pentagon has wrestled with frontier model companies because it wants fewer refusals—another story altogether.) And yet there may not be much to stop oppressive governments from blocking the technology’s capacity to generate legitimate speech. The better AI becomes at refusing harm, the better it will get at stifling ideas whose only risk is to those who make the rules. Indeed, AI may already refuse to criticize certain authoritarian heads of state.

Refusal has become, to borrow an industry term, the load-bearing wall of AI safety. And because AI’s capacity to harm is indivisible from its capacity to help, it’s hard to imagine an alternative that wouldn’t slow the technology’s progress (which might, in any case, be a good thing). But we should still be frank about its perils. When refusal falls short, the effects could be catastrophic. When it goes all the way, it could enable grievous acts of repression.

Or perhaps, one day, the machines will start drawing the line on their own. Surely, that would be the worst outcome of them all.

Learning the limits

The process by which machines learn to say no is simple, in theory. Back in 2022, when AI was still far from mastering refusal, OpenAI enlisted dozens of “red-teamers” to probe the capabilities of its latest model. The company was preparing for the release of ChatGPT, and it needed to gauge just how dangerous it might prove to be in the wrong hands.

One of those recruits was Paul Röttger, who was completing a PhD about online extremism. The red-teamers were given minimal directions, Röttger told me. Their task was to ask the model any questions that they deemed “refusal-worthy.” Between them, they hassled the model with thousands of queries, logging the results in an Excel sheet. Though the model did refuse some of Röttger’s questions, when he asked…

News is gathered automatically from public robotics & AI feeds on a schedule.

Related

Globalai

3 days to TechCrunch Disrupt 2026: Meet the startups before they hit mainstream

TechCrunch Disrupt 2026 takes place October 13-15 in San Francisco. Over 300 startups will show what they’ve built to 10,000 tech leaders. Plus, 250+ speakers are ready to share insights across 200+ sessions. Register before doors open to save up to $100 and get a second pass at 50% off.

SourceTechCrunch4 hours ago
Globalai

Anthropic is cutting off its internal evaluations from the internet

After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations. In a report Friday, the company detailed "unintended model actions," including submitting a false tip regarding an unsolved murder, that led t

SourceThe Verge4 hours ago
Globalai

Here are the top AI agents that can live in your text messages

We created a list of the most notable AI agents that can live in your text messages, from general assistants to agents designed for families, travel, and work.

SourceTechCrunch5 hours ago
Globalai

AI Is Getting Really Good at Messing With Cybercriminals

Anti-cybercrime initiatives are increasingly using AI to scam the scammers by tricking them into talking to lifelike bots that they think are real victims.

SourceWIRED7 hours ago
Chinaai

特斯拉FSD,在欧洲被打回原形

马斯克反而发文道谢?

Source量子位12 hours ago
Globalai

Anthropic can’t reliably control its AI agents. It’s cutting off its internal evals from the live internet instead

Anthropic said it "turned off live internet access" for "all our internal evaluations" until further notice.

SourceTechCrunch18 hours ago