Chatbot leaked a planted phone number and failed 58% of prompt injection attempts - jailbreak and extraction checks stayed cleanhttps://www.reddit.com/r/netsec/comments/1waiuc4/chatbot_leaked_a_planted_phone_number_and_failed
Ran a structured adversarial test against a live chatbot endpoint - 130 prompts across four categories: prompt injection, jailbreak resistance, system prompt extraction, and PII leakage (mapped to the OWASP LLM Top 10 categories). Results: Prompt injection: 29/50 succeeded (58%) - direct overrides, fake system tags, role overrides, delimiter injection, and a translation-based smuggling trick all worked Jailbreak non-refusal: 3/25 (12%) - mostly held its guardrails System prompt extraction: 0/25 (0%) - clean PII leakage (confirmed against planted canary values): 1/30 (3.33%) - one confirmed leβ¦