Llama Guard logo

Llama GuardOpen LLM-based safeguard for classifying unsafe content in human-AI conversations.

4.6 (5)

Overview

Llama Guard is a safety classifier built on top of Meta's Llama models, designed to evaluate both user prompts and model responses for potentially harmful content. It outputs a safety label along with the specific policy categories that were violated, making it useful as a guardrail layer around chatbots and other generative AI systems. The model is trained against a configurable taxonomy covering categories such as violence, sexual content, hate, self-harm, and criminal advice. Because the taxonomy is provided in the prompt itself, developers can adapt or extend the policy without retraining, tailoring moderation to their specific application or jurisdiction. Distributed with open weights, Llama Guard can be self-hosted alongside an LLM pipeline to filter inputs and outputs in real time, offering an alternative to closed moderation APIs for teams that need transparency, customization, or on-premise deployment.

Key features

  • LLM-based input and output moderation
  • Multi-category harm classification
  • Prompt-configurable policy taxonomy
  • Open-source weights from Meta
  • Compatible with Llama and other LLM stacks
  • Returns safe/unsafe label with violated categories

Pricing

Model
Freemium
Rating
4.6 / 5 (5)

Use cases

Chatbot input and output moderation

Wrap a production chatbot with Llama Guard to screen user prompts and model responses, blocking unsafe content before it reaches end users.

Custom policy enforcement

Adapt the prompt-based taxonomy to match an application's specific policies or jurisdictional requirements without retraining the safety model.

Self-hosted compliance layer

Deploy open weights on-premises to audit and moderate LLM traffic in regulated environments where data cannot leave internal infrastructure.

Red-teaming and dataset filtering

Use Llama Guard to label conversation datasets for unsafe categories, supporting safety evaluations, fine-tuning data curation, and red-team analysis.

Pros & Cons

Pros

  • Open weights enable self-hosting and auditing
  • Customizable safety taxonomy via prompt
  • Classifies both user inputs and model outputs
  • Integrates easily into existing LLM pipelines

Cons

  • Requires GPU resources to run efficiently
  • May produce false positives or miss nuanced harms
  • Setup and tuning expertise needed
  • English-centric performance

Battle record

Across 4 battles in the Pantheon.

1
1st
0
2nd
0
3rd

Last 4 battles

Reviews

4.6

Average from 5 ratings.

5
3
4
2
3
0
2
0
1
0

Sign in to leave a review.

Tomáš Novák

Tomáš Novák

Apr 4, 2026

Use it every day

Honestly didn't expect to like it this much. Compatible with Llama and other LLM stacks is exactly what I needed, and integrates easily into existing LLM pipelines. but I reach for it almost every day now and it just clicks.

IB

Ingrid Bauer

Mar 24, 2026

Solid for our team

We rolled this out across the team last quarter and open weights enable self-hosting and auditing. LLM-based input and output moderation fits neatly into how we already work, and compatible with Llama and other LLM stacks removed a step we used to do by hand. but it has held up under daily use.

TA

Tariq Aziz

Feb 22, 2026

Use it every day

Honestly didn't expect to like it this much. Compatible with Llama and other LLM stacks is exactly what I needed, and open weights enable self-hosting and auditing. but I reach for it almost every day now and it just clicks.

Daniel Schmidt

Daniel Schmidt

Sep 6, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is compatible with Llama and other LLM stacks — handled better than most — and open weights enable self-hosting and auditing. Requires GPU resources to run efficiently is my one real gripe. Worth the time if this is your use case.

Aaliyah Johnson

Aaliyah Johnson

Jun 17, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: lLM-based input and output moderation and integrates easily into existing LLM pipelines. Where it lags: english-centric performance. On balance the feature set — especially lLM-based input and output moderation — justifies the 4 stars for our use case.

Q&A

Is Llama Guard compatible with other LLM stacks?

Yes, Llama Guard is compatible with Llama and other LLM stacks, and can be self-hosted alongside an LLM pipeline.

Asked by Jarrah Whitlock · Sep 19, 2025

What are the system requirements to run Llama Guard?

Llama Guard requires GPU resources to run efficiently.

Asked by Sven Bergqvist · Aug 12, 2025

Can I customize the safety taxonomy?

Yes, the taxonomy is provided in the prompt itself, allowing developers to adapt or extend the policy without retraining, tailoring moderation to their specific application or jurisdiction.

Asked by Chioma Nwosu · Aug 11, 2025

What is Llama Guard used for?

Llama Guard is used for classifying unsafe content in human-AI conversations, evaluating both user prompts and model responses for potentially harmful content.

Asked by Faisal Rahman · Jun 19, 2025

Ask a question

Predictive Analytics alternatives