πŸ” Search
Sign in to post
Chatbot leaked a planted phone number and failed 58% of prompt injection attempts - jailbreak and extraction checks stayed cleanhttps://www.reddit.com/r/netsec/comments/1waiuc4/chatbot_leaked_a_planted_phone_number_and_failed

Ran a structured adversarial test against a live chatbot endpoint - 130 prompts across four categories: prompt injection, jailbreak resistance, system prompt extraction, and PII leakage (mapped to the OWASP LLM Top 10 categories). Results: Prompt injection: 29/50 succeeded (58%) - direct overrides, fake system tags, role overrides, delimiter injection, and a translation-based smuggling trick all worked Jailbreak non-refusal: 3/25 (12%) - mostly held its guardrails System prompt extraction: 0/25 (0%) - clean PII leakage (confirmed against planted canary values): 1/30 (3.33%) - one confirmed le…

Hacking AI customer service agents (Bug Bounty Village DEF CON 34)https://www.reddit.com/r/netsec/comments/1wae5zc/hacking_ai_customer_service_agents_bug_bounty

At Bug Bounty Village during DEF CON 34, Inti De Ceukelaire delivered a talk on how attackers can abuse today's AI agents in ways most defenders haven't thought about yet, from tricking agents into spilling secrets to forcing them to carry out unauthorized actions on behalf of the victim. This resulted in over $50,000+ in bounties in just a few weekends, without actually poking the target with Burp Suite or any automated scanners. submitted by /u/qwerty0x41 [link] [comments]

0trust.social media

Loading your media...

Pick a GIF β€” Giphy

Loading GIFs...