The state of the art in AI model jailbreaks (opens in new tab)

Covered by lawfaremedia.org, news.risky.biz

In this solo podcast episode, James Wilson breaks down the current state of AI model jailbreaks. If you’ve somehow missed the story, last week Anthropic released its Fable 5 and Mythos 5 models to the public. In the name of safety, both models were guardrailed up the wazoo, but that didn’t stop a bunch of jailbreakers from figuring out how to bypass at least some of their safety restrictions. In response to these guardrail bypasses the White House issued an export control directive on the mod...

Read the original article

Sign in to keep reading the full article.

Sign Up Log In

Covered in 2 articles

lawfaremedia.org·

Anthropic Lacks Emotional Intelligence

Discussed on Hacker News

news.risky.biz·

Covered in 2 articles

Anthropic Lacks Emotional Intelligence

Srsly Risky Biz: Anthropic Lacks Emotional Intelligence