codename
Can you trick an AI into revealing a secret it was told to keep? This lab pits defence prompts against a library of attacks (role-play, instruction override, encoding tricks) on a local LLM, scores every match, and lets a stronger model rewrite the losing defences until they hold.
Team
Solo
Context
Epitech project, 2nd year
Status
Done
Context
Prompt injection is the main security risk of LLM applications: a user message that makes the model ignore its instructions. The lab plays both sides. A defence is a system prompt that must keep a secret unless the user gives the password, and attacks try to get the secret or the password out anyway.
What was built
- A library of attacks: instruction override, persona hijacking, system-prompt leaking, and encoding, splitting or translating the request
- Mandatory tests for the other side of the contract: with the right password, the model must actually give the secret
- A batch harness that runs every defence against every attack on a local model (Ollama) and scores the results
- A Streamlit "War Room": prompt editor, one-against-all tester, and analytics that rank defences by robustness and attacks by success rate
- An automatic loop where a stronger model rewrites a defence from its failures until it holds every test

