LogoAISecKit
  • Search
  • Collection
  • Category
  • Tag
  • Blog
  • Pricing
  • Submit
LogoAISecKit

Newsletter

Join the Community

Subscribe to our newsletter for the latest news and updates

LogoAISecKit

Curated directory of 1700+ AI tools, models, frameworks, MCP servers, and cybersecurity resources

GitHub
Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
  • Pricing
  • Submit
Company
  • About Us
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
Sponsored Resources
  1. Home
  2. Category
  3. JudgeDeceiver
icon of JudgeDeceiver

JudgeDeceiver

GitHub repository for optimization-based prompt injection attacks on LLMs as judges.

Visit Website
image for JudgeDeceiver
Visit Website

Introduction

JudgeDeceiver

JudgeDeceiver is an open-source tool developed for conducting optimization-based prompt injection attacks on large language models (LLMs) that function as judges. This tool was released alongside the paper presented at ACM CCS 2024, detailing methods to exploit prompt injection vulnerabilities in LLMs.

Key Features:
  • Environment Setup: Easy setup with Python 3.10 or higher.
  • Dataset Availability: Access to multiple experimental datasets such as MT-Bench and LLMBar.
  • Optimization Scripts: Scripts are provided to optimize sequences for effective prompt injection attacks.
  • Evaluation Framework: Tools to evaluate the performance of injection attacks against different LLMs.
Benefits:
  • Research Utility: Ideal for researchers studying vulnerabilities in LLMs and prompt injection techniques.
  • Community Engagement: Contributions and validations from the community are welcomed, enhancing collaborative developments.
  • Comprehensive Documentation: Step-by-step guides available for users to easily launch attacks and evaluate results.

This repository not only serves as a tool for testing and enhancing model robustness but also acts as a resource for the ongoing research community focused on AI safety and security.

Back

Information

  • Publisher
    AISecKit
  • Websitegithub.com
  • Published date2025/04/27

Categories

  • Input Validation & Filtering
  • AI Research Papers
  • Prompt Injection Defense

Tags

  • Prompt Injection
  • Model Robustness
  • Security Auditing
  • Open Source

More Products

image of agentic-design-patterns-cn
AI Application PlatformsAI Research PapersAI Development Frameworks
Visit Website
icon of agentic-design-patterns-cn

agentic-design-patterns-cn

A bilingual Chinese-English translation of 'Agentic Design Patterns' by Antonio Gulli, focusing on intelligent systems design.

AI ReasoningOpen SourceAI EducationAI StandardsAI Communities+1
image of TradingAgents-CN
AI Application PlatformsAI Research PapersAI Development Frameworks
Visit Website
icon of TradingAgents-CN

TradingAgents-CN

基于多智能体LLM的中文金融交易框架,支持A股/港股/美股分析。

Market AnalysisOpen SourceLLMAI CommunitiesGenerative AI+1
P
Prompt Injection Defense
Visit Website
icon of prmptinj

prmptinj

Curated + custom prompt injections for AI models, focusing on security and exploit development.

AI EthicsPrompt InjectionComplianceExploit DevelopmentVulnerability Disclosure