πŸ” Search
Sign in to post
Chatbot leaked a planted phone number and failed 58% of prompt injection attempts - jailbreak and extraction checks stayed cleanhttps://www.reddit.com/r/netsec/comments/1waiuc4/chatbot_leaked_a_planted_phone_number_and_failed

Ran a structured adversarial test against a live chatbot endpoint - 130 prompts across four categories: prompt injection, jailbreak resistance, system prompt extraction, and PII leakage (mapped to the OWASP LLM Top 10 categories). Results: Prompt injection: 29/50 succeeded (58%) - direct overrides, fake system tags, role overrides, delimiter injection, and a translation-based smuggling trick all worked Jailbreak non-refusal: 3/25 (12%) - mostly held its guardrails System prompt extraction: 0/25 (0%) - clean PII leakage (confirmed against planted canary values): 1/30 (3.33%) - one confirmed le…

0trust.social media

Loading your media...

Pick a GIF β€” Giphy

Loading GIFs...