ModelsAgree
← All leaderboards

garak

What ChatGPT, Claude, Gemini & Grok actually say · August 2026

Visit garak.ai

The verdict

garak appears in 1 AI-ranked category — best position #2 for ai red teaming and llm security testing tool.

Positioning brief — for the garak team

Why the models put garak at #2 for ai red teaming and llm security testing tool

  • Broad actively maintained probe library Claude · Grok · Gemini · GPTa huge, actively maintained probe library
  • Rapid automated baseline vulnerability scanning Claude · Grok · Gemini · GPTrapid, automated baseline vulnerability scanning
  • Free open-source and model-agnostic Claude · GPTfree, model-agnostic, and NVIDIA-backed
  • Extensible with broad model support Claude · Grok · GPTreproducible testing, broad model support

What the models credit Promptfoo (#1) with — and don’t credit garak

  • Agent and RAG testing GPT · Grokagent/RAG testing
  • Configurable multi-turn strategies GPTconfigurable multi-turn strategies
  • CI/CD continuous regression testing GPT · Gemini · Claude · Grokprovides continuous regression testing

What would move the rank — the models’ fix lines, unified

  • Add adaptive multi-turn agentic attacks GPT · Claude · Gemini · Grokadaptive multi-turn/agentic attacks
  • Expand app-specific RAG exploitation GPT · Claude · Grokstateful end-to-end agent, tool, and RAG exploitation
  • Improve CI/CD integration and reporting Claude · Groktriage/reporting is DIY

Restructured from verbatim model output · nothing invented · every quote machine-verified

GPT #5Claude #1Gemini #2Grok #1

The de facto open-source LLM vulnerability scanner — a huge, actively maintained probe library covering jailbreaks, prompt injection, data/PII leakage, toxicity, encoding attacks, and package hallucination, runnable against virtually any model (HuggingFace, OpenAI, local, custom REST) with a single CLI command; free, model-agnostic, and NVIDIA-backed, it's the closest thing to "nmap for LLMs" and the fastest way for a security engineer to get real coverage of known attack classes. Ranked #1 on the assumption the typical practitioner is a security/ML engineer wanting broad, off-the-shelf coverage rather than a managed service.

Grok Broadest probe library (120+ categories covering prompt injection, jailbreaks, hallucination, data leakage, toxicity, encoding attacks) with generators/detectors for systematic scanning of models and dialog systems; battle-tested, extensible, high public evidence of use in research/security workflows; strong for model-level and basic app testing.

Gemini The standard "nmap of LLMs" that provides rapid, automated baseline vulnerability scanning across hundreds of pre-built jailbreak, toxicity, and data leakage probes.

GPT Mature open-source LLM vulnerability scanner with a large probe ecosystem, reproducible testing, broad model support, and strong value for baseline security assessments

Where garak falls short, per the models

  • GPT Expand beyond model-centric probes into stateful end-to-end agent, tool, and RAG exploitation
  • Claude It's a static probe scanner — strong on cataloged attack patterns but weaker on adaptive multi-turn/agentic attacks and app-specific business-logic flaws, and triage/reporting is DIY, so it's not for teams wanting a polished dashboard or continuous managed testing.
  • Gemini Lacks the ability to simulate complex multi-turn agent interactions and relies heavily on static detectors that can miss subtle exploits.
  • Grok Primarily stateless/single-turn focused model scanning (less native multi-turn/agentic depth without heavy customization); not ideal for full production app pipelines or teams needing CI/CD regression without extra integration.

Poll history — On this board 2 of 2 polls since Jul 12 · now #1

#5#1

What changed in the models’ minds

GeminiJul 12Jul 13 poll

  • Newtoxicity probestoxicity
  • Newstatic detectors miss subtle exploitsrelies heavily on static detectors that can miss subtle exploits
  • Droppedmulti-modal inputs

Top alternatives per the models: Promptfoo · PyRIT · Mindgard · Giskard

Head-to-head — how the models call it

Watch garak

Boards re-poll weekly and the models change their minds. One short email only when garak's standing moves — a rank change, a rival overtaking, or new reasoning from the models. Nothing otherwise.

Embed your ranking badge

garak ranks #2 for best ai red teaming and llm security testing tool by AI-model consensus. Put the badge in your README, docs or site — it updates automatically as the models re-rank.

garak — ranked #2 for Best AI red teaming and LLM security testing tool by AI models on ModelsAgree
Markdown (README)
[![garak — ranked #2 for Best AI red teaming and LLM security testing tool by AI models on ModelsAgree](https://modelsagree.com/badge/garak.svg)](https://modelsagree.com/best/best-ai-red-teaming-tool?utm_source=badge&utm_medium=embed&utm_campaign=badge-garak)
HTML
<a href="https://modelsagree.com/best/best-ai-red-teaming-tool?utm_source=badge&utm_medium=embed&utm_campaign=badge-garak"><img src="https://modelsagree.com/badge/garak.svg" alt="garak — ranked #2 for Best AI red teaming and LLM security testing tool by AI models on ModelsAgree" height="28"></a>

Rankings are computed from what the models answer, re-polled on demand · raw reasoning shown verbatim · methodology