Skip to main content
Scour
Discover
Docs
Login
Sign Up
You are offline. Trying to reconnect...
Copied to clipboard
Unable to share or copy to clipboard
The AI Security Institute (AISI)
aisi.gov.uk
AI Security Institute
·
17h
17 hours ago
Safety case template for frontier AI: A cyber inability argument
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Safety case template for frontier AI: A cyber inability argument
AI Security Institute
·
17h
17 hours ago
Early Insights from Developing Question-Answer Evaluations for Frontier AI | AISI Work
Covers
4 stories
See all stories this covers
including
Language models are few-shot learners (2020)
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Early Insights from Developing Question-Answer Evaluations for Frontier AI | AISI Work
AI Security Institute
·
17h
17 hours ago
Adversarial machine learning: A taxonomy and terminology of attacks and mitigations
Covers
Adversarial ML: A Taxonomy and Terminology of Attacks / Mitigations
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Adversarial machine learning: A taxonomy and terminology of attacks and mitigations
AI Security Institute
·
17h
17 hours ago
The Inspect Sandboxing Toolkit: Scalable and secure AI agent evaluations | AISI Work
Covers
3 stories
See all stories this covers
including
Featured Research
Covered by
pub.towardsai.net
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for The Inspect Sandboxing Toolkit: Scalable and secure AI agent evaluations | AISI Work
AI Security Institute
·
17h
17 hours ago
Early lessons from evaluating frontier AI systems | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Early lessons from evaluating frontier AI systems | AISI Work
AI Security Institute
·
17h
17 hours ago
Our approach to tackling AI-generated child sexual abuse material | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Our approach to tackling AI-generated child sexual abuse material | AISI Work
AI Security Institute
·
17h
17 hours ago
LLM judges on trial: A new statistical framework to assess autograders | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for LLM judges on trial: A new statistical framework to assess autograders | AISI Work
AI Security Institute
·
17h
17 hours ago
Ask Don't Tell: Reducing Sycophancy in Large Language Models | AISI Work
Covers
2 stories
See all stories this covers
including
AI chatbots are becoming "sycophants" to drive engagement, a new study of 11 leading models finds
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Ask Don't Tell: Reducing Sycophancy in Large Language Models | AISI Work
AI Security Institute
·
17h
17 hours ago
STACK: Adversarial attacks on LLM safeguard pipelines
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for STACK: Adversarial attacks on LLM safeguard pipelines
AI Security Institute
·
17h
17 hours ago
An alignment safety case sketch based on debate
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for An alignment safety case sketch based on debate
AI Security Institute
·
17h
17 hours ago
Introducing ControlArena: A library for running AI control experiments | AISI Work
Covers
Inspect
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Introducing ControlArena: A library for running AI control experiments | AISI Work
AI Security Institute
·
17h
17 hours ago
Inspect Cyber: A New Standard for Agentic Cyber Evaluations | AISI Work
Covers
Inspect
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Inspect Cyber: A New Standard for Agentic Cyber Evaluations | AISI Work
AI Security Institute
·
17h
17 hours ago
Infusion: Shaping model behaviour by editing training data via influence functions
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Infusion: Shaping model behaviour by editing training data via influence functions
AI Security Institute
·
17h
17 hours ago
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
AI Security Institute
·
17h
17 hours ago
Breaking agent backbones: Evaluating the security of backbone LLMs in AI agents
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Breaking agent backbones: Evaluating the security of backbone LLMs in AI agents
AI Security Institute
·
17h
17 hours ago
How do AI models persuade? Exploring the levers of AI-enabled persuasion through large-scale experiments | AISI Work
Covers
The levers of political persuasion with conversational artificial intelligence
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for How do AI models persuade? Exploring the levers of AI-enabled persuasion through large-scale experiments | AISI Work
AI Security Institute
·
17h
17 hours ago
New updates to the AISI Challenge Fund | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for New updates to the AISI Challenge Fund | AISI Work
AI Security Institute
·
17h
17 hours ago
A sketch of an AI control safety case
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for A sketch of an AI control safety case
AI Security Institute
·
17h
17 hours ago
Boundary Point Jailbreaking: A new way to break the strongest AI defences | AISI Work
Covers
3 stories
See all stories this covers
including
Constitutional Classifiers: Defending against universal jailbreaks
Covered by
LessWrong
,
alignment.anthropic.com
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Boundary Point Jailbreaking: A new way to break the strongest AI defences | AISI Work
AI Security Institute
·
17h
17 hours ago
Stress-testing asynchronous monitoring of AI coding agents | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Stress-testing asynchronous monitoring of AI coding agents | AISI Work
« Page 1
·
Page 3 »
Log in to enable infinite scrolling
Keyboard Shortcuts
Navigation
Next / previous post
j
/
k
Open post
o
or
Enter
Preview post
v
Post Actions
Love post
a
Like post
l
Dislike post
d
Undo reaction
u
Save / unsave
s
Recommendations
Add interest / feed
Enter
Not interested
x
Go to
Home
g
h
Interests
g
i
Feeds
g
f
Likes
g
l
History
g
y
Changelog
g
c
Settings
g
s
Discover
g
b
Search
/
Pagination
Next page
n
Previous page
p
General
Show this help
?
Submit feedback
!
Close modal / unfocus
Esc
Press
?
anytime to show this help
Like
Save
Not for me
Report