Skip to main content
Scour
Discover
Docs
Login
Sign Up
You are offline. Trying to reconnect...
Copied to clipboard
Unable to share or copy to clipboard
The AI Security Institute (AISI)
aisi.gov.uk
AI Security Institute
·
16h
16 hours ago
Navigating the uncharted: Building societal resilience to frontier AI | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Navigating the uncharted: Building societal resilience to frontier AI | AISI Work
AI Security Institute
·
16h
16 hours ago
Open technical problems in open-weight AI model risk management
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Open technical problems in open-weight AI model risk management
AI Security Institute
·
16h
16 hours ago
Evaluating whether AI models would sabotage AI safety research | AISI Work
Covers
2 stories
See all stories this covers
including
Anthropic's Autonomous AI Agents Outperform Human Researchers on Weak-to-Strong Supervision
Covered by
LessWrong
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Evaluating whether AI models would sabotage AI safety research | AISI Work
AI Security Institute
·
16h
16 hours ago
How are AI agents used? Evidence from 177,000 MCP tools
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for How are AI agents used? Evidence from 177,000 MCP tools
AI Security Institute
·
16h
16 hours ago
Should AI systems behave like people? | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Should AI systems behave like people? | AISI Work
AI Security Institute
·
16h
16 hours ago
Lessons from a chimp: AI "scheming" and the quest for ape language
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Lessons from a chimp: AI "scheming" and the quest for ape language
AI Security Institute
·
16h
16 hours ago
Boundary Point Jailbreaking of Black-Box LLMs
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Boundary Point Jailbreaking of Black-Box LLMs
AI Security Institute
·
16h
16 hours ago
Why human-AI relationships need socioaffective alignment
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Why human-AI relationships need socioaffective alignment
AI Security Institute
·
16h
16 hours ago
Mapping the limitations of current AI systems | AISI Work
Covers
6 stories
See all stories this covers
including
Measuring AI Ability to Complete Long Tasks
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Mapping the limitations of current AI systems | AISI Work
AI Security Institute
·
16h
16 hours ago
Model tampering attacks enable more rigorous evaluations of LLM capabilities
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Model tampering attacks enable more rigorous evaluations of LLM capabilities
AI Security Institute
·
16h
16 hours ago
Releasing AISI’s Engineering Playbook | AISI Work
Covers
4 stories
See all stories this covers
including
International AI Safety Report 2026
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Releasing AISI’s Engineering Playbook | AISI Work
AI Security Institute
·
16h
16 hours ago
International evaluation best practice and open questions in AI measurement | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for International evaluation best practice and open questions in AI measurement | AISI Work
AI Security Institute
·
16h
16 hours ago
Managing risks from increasingly capable open-weight AI systems | AISI Work
Covers
4 stories
See all stories this covers
including
OpenAI/GPT-OSS-120B · Hugging Face
Covered by
War on the Rocks
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Managing risks from increasingly capable open-weight AI systems | AISI Work
AI Security Institute
·
16h
16 hours ago
Principles for safeguard evaluation | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Principles for safeguard evaluation | AISI Work
AI Security Institute
·
16h
16 hours ago
Pre-deployment evaluation of Anthropic’s upgraded Claude 3.5 Sonnet | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Pre-deployment evaluation of Anthropic’s upgraded Claude 3.5 Sonnet | AISI Work
AI Security Institute
·
16h
16 hours ago
Automated alignment is harder than you think
Covers
Automated alignment is harder than you think
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Automated alignment is harder than you think
AI Security Institute
·
16h
16 hours ago
Ask don't tell: Reducing sycophancy in large language models
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Ask don't tell: Reducing sycophancy in large language models
AI Security Institute
·
16h
16 hours ago
RepliBench: measuring autonomous replication capabilities in AI systems | AISI Work
Covers
4 stories
See all stories this covers
including
hypothetical tendency for sufficiently intelligent agents to pursue unbounded instrumental goals such as self-preservation and resource acquisition
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for RepliBench: measuring autonomous replication capabilities in AI systems | AISI Work
AI Security Institute
·
16h
16 hours ago
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
AI Security Institute
·
16h
16 hours ago
Prefill Awareness in Large Language Models
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Prefill Awareness in Large Language Models
Page 2 »
Log in to enable infinite scrolling
Keyboard Shortcuts
Navigation
Next / previous post
j
/
k
Open post
o
or
Enter
Preview post
v
Post Actions
Love post
a
Like post
l
Dislike post
d
Undo reaction
u
Save / unsave
s
Recommendations
Add interest / feed
Enter
Not interested
x
Go to
Home
g
h
Interests
g
i
Feeds
g
f
Likes
g
l
History
g
y
Changelog
g
c
Settings
g
s
Discover
g
b
Search
/
Pagination
Next page
n
Previous page
p
General
Show this help
?
Submit feedback
!
Close modal / unfocus
Esc
Press
?
anytime to show this help
Like
Save
Not for me
Report