Show HN: Semantic Overlays – an NX bit for LLM prompt injection (live demo) https://ift.tt/iVNq9cm
Show HN: Semantic Overlays – an NX bit for LLM prompt injection (live demo) I've built a new method for steering LLMs called Semantic Overlays, small trained adapters on a frozen model which change how it perceives a piece of its context. The most readily applicable usage is to mitigate prompt injection, and it lets us take a very-injectable Qwen-3.5-9B to SOTA scores on all the prompt injection benchmarks I could find. (They are only blackbox attacks, but I did NOT train on anything like them — whitebox attacks are out of scope for this paper) I'm excited for you to play with the tech — see if YOU can break it! (let me know if you can) Paper at https://ift.tt/lOJMxfq if you want to read more about it, code at https://ift.tt/lzQi5kc , adapters at https://ift.tt/xBkVXOe... Also https://ift.tt/7zmOEcl if you wanna watch a little video I made! https://semantic-overlays.vercel.app/ September 2, 2026 at 12:40AM
Komentar
Posting Komentar