A Theory of Prompt Injection (and why you should study roles) (opens in new tab)

Covers 2 stories including Playwright MCP Server – Snapshot based – faster and more reliable than images

SummaryWe've been building a theory of how prompt injections work under the hood.We show it comes down to how LLMs perceive roles (the humble chat template tags).We use this theory to create new attacks, explain some weird mech interp results, and predict when attacks work.We also advocate for a new subfield focused on the science of roles, and sketch some unexplored new research problems.Work supported by CBAI and Cosmos. Another version of this post (with more inline colors) is here, and fu...

Read the original article