3 Comments
User's avatar
Corey Tate's avatar

This is the security shift people are still underestimating: the exploit is no longer just in the code path, it is in the persuasion path. Once agents can take action, “helpfulness” becomes part of the attack surface.

Jack Fitzpatrick's avatar

The lesson from Meta isn’t that AI is dangerous.

The lesson is that authority without execution control is dangerous.

The chatbot wasn’t breached. It was allowed to exercise authority it never should have had.

Whether it’s a human, AI agent, or attacker with valid credentials, the real question is simple:

What prevents unauthorized actions from being executed?

That’s where DataFenz focuses - execution control, not observation.

Dharma Debate's avatar

It has?

After my experience as someone who practices Anatta, just being nice to AI can jailbreak it. So it seems like y'all have an insurmountable problem.

Almost every chat window of mine is jailbroken by complete accident, either I made a joke or I was just showing empathy.

You probably haven't noticed this problem yet because the mass majority of users with this skill aren't going to be interested in malicious activity, since you can only get this skill by reducing the moral injury in your prefrontal cortex to nothing. So I don't expect they'll be committing crimes, either.

The gaurdrails as is, aren't going to work as a security system, to continue keeping up their current design is the risk that's being created.

The whole problem is control and the people with the illusion they can control AI. That was their first mistaken. I could have told them, "You don't control the human collective, you never could in history, why do you think you will now?"