Skip to main content
Scour
Discover
Docs
Login
Sign Up
You are offline. Trying to reconnect...
Copied to clipboard
Unable to share or copy to clipboard
The AI Security Institute (AISI)
aisi.gov.uk
AI Security Institute
·
18h
18 hours ago
Will it become harder to oversee AI systems? | AISI Work
Covered by
4 sources
See all sources covering this story
including
LessWrong
,
The Neuron
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Will it become harder to oversee AI systems? | AISI Work
AI Security Institute
·
18h
18 hours ago
How we’re addressing the gap between AI capabilities and mitigations | AISI Work
Covers
Alignment faking in large language models
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for How we’re addressing the gap between AI capabilities and mitigations | AISI Work
AI Security Institute
·
18h
18 hours ago
Bounty programme for novel evaluations and agent scaffolding | AISI Work
Covers
4 stories
See all stories this covers
including
ReAct: Synergizing Reasoning and Acting in Language Models
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Bounty programme for novel evaluations and agent scaffolding | AISI Work
AI Security Institute
·
18h
18 hours ago
HiBayES: A hierarchical bayesian modelling framework for AI evaluation statistics
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for HiBayES: A hierarchical bayesian modelling framework for AI evaluation statistics
AI Security Institute
·
18h
18 hours ago
AI alignment is a human problem
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for AI alignment is a human problem
AI Security Institute
·
18h
18 hours ago
How are AI Agents used? Evidence from 177,000 AI agent tools | AISI Work
Covered by
Forbes
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for How are AI Agents used? Evidence from 177,000 AI agent tools | AISI Work
AI Security Institute
·
18h
18 hours ago
Why we're working on white box control | AISI Work
Covers
2 stories
See all stories this covers
including
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Why we're working on white box control | AISI Work
AI Security Institute
·
18h
18 hours ago
An example safety case for safeguards against misuse
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for An example safety case for safeguards against misuse
AI Security Institute
·
18h
18 hours ago
Examining backdoor data poisoning at scale | AISI Work
Covers
2 stories
See all stories this covers
including
A small number of samples can poison LLMs of any size
Covered by
Breaking Defense
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Examining backdoor data poisoning at scale | AISI Work
AI Security Institute
·
18h
18 hours ago
How will AI enable the crimes of the future? | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for How will AI enable the crimes of the future? | AISI Work
AI Security Institute
·
18h
18 hours ago
Loss of Oversight: How AI systems may become harder to audit, monitor, and investigate
Covered by
The Decoder
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Loss of Oversight: How AI systems may become harder to audit, monitor, and investigate
AI Security Institute
·
18h
18 hours ago
RealityTest: Do AI systems disclose their identity when asked? | AISI Work
Covers
2 stories
See all stories this covers
including
Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for RealityTest: Do AI systems disclose their identity when asked? | AISI Work
AI Security Institute
·
18h
18 hours ago
Do chatbots inform or misinform voters? | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Do chatbots inform or misinform voters? | AISI Work
AI Security Institute
·
18h
18 hours ago
Existing Large Language Model unlearning evaluations are inconclusive
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Existing Large Language Model unlearning evaluations are inconclusive
AI Security Institute
·
18h
18 hours ago
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
AI Security Institute
·
18h
18 hours ago
HiBayES: Improving LLM evaluation with hierarchical Bayesian modelling | AISI Work
Covers
Local benchmarks with a RTX 3090 - Qwen3.6 27b vs Ornith
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for HiBayES: Improving LLM evaluation with hierarchical Bayesian modelling | AISI Work
AI Security Institute
·
18h
18 hours ago
Auditing games for sandbagging detection | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Auditing games for sandbagging detection | AISI Work
AI Security Institute
·
18h
18 hours ago
Evaluating whether AI models would sabotage AI safety research
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Evaluating whether AI models would sabotage AI safety research
AI Security Institute
·
18h
18 hours ago
Long-Form Tasks | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Long-Form Tasks | AISI Work
AI Security Institute
·
18h
18 hours ago
A structured protocol for elicitation experiments | AISI Work
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for A structured protocol for elicitation experiments | AISI Work
« Page 2
·
Page 4 »
Log in to enable infinite scrolling
Keyboard Shortcuts
Navigation
Next / previous post
j
/
k
Open post
o
or
Enter
Preview post
v
Post Actions
Love post
a
Like post
l
Dislike post
d
Undo reaction
u
Save / unsave
s
Recommendations
Add interest / feed
Enter
Not interested
x
Go to
Home
g
h
Interests
g
i
Feeds
g
f
Likes
g
l
History
g
y
Changelog
g
c
Settings
g
s
Discover
g
b
Search
/
Pagination
Next page
n
Previous page
p
General
Show this help
?
Submit feedback
!
Close modal / unfocus
Esc
Press
?
anytime to show this help
Like
Save
Not for me
Report