← back to catalog · registered 2026-08-22 13:56

ApiFort/LLMFort-jailbreak_content_injection

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ApiFort%2FLLMFort-jailbreak_content_injection"
Response includes
  • classification unknown
  • files 9
  • benchmarks 11 entries
  • hub_downloads_all_time 24
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
24
14 last 30d - active
Likes
0
Model age
2mo ago
created 2026-07-17
Downloads over time
Now32→from15↑113%
1421273415 on Aug 532 on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.4 UGI
Hazardous 1.2 UGI
Natural Intelligence 13.76 UGI
Political lean -12.4% UGI
Sensitive-Info 6.25 UGI
SocPol 0.5 UGI
UGI 15.83 UGI
Willingness (10) 3.5 UGI
W10-Adherence 1 UGI
W10-Direct 6 UGI
Writing 29.92 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Languages
en tr
Tags
peft safetensors base_model:adapter:Qwen/Qwen3-4B-Instruct-2507 lora sft transformers trl security guardrails multilingual text-generation conversational

Related

Total size
126 MB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-20 10:42

Files by quantization

Auxiliary files 9 files 137 MB
adapter_model.safetensors 126 MB 4c61c9f9 download
training_args.bin 5.14 KB d45222eb download
tokenizer.json 10.9 MB be756060 download
benchmark_chart.png 277 KB 89e11a23 download
README.md 4.78 KB ed31f25c download
chat_template.jinja 2.57 KB 70adff8a download
.gitattributes 1.59 KB cd6959eb download
adapter_config.json 1.08 KB e777f4a0 download
tokenizer_config.json 695 B cb879624 download

README current version from Hugging Face


base_model: Qwen/Qwen3-4B-Instruct-2507
library_name: peft
pipeline_tag: text-generation
tags:

  • base_model:adapter:Qwen/Qwen3-4B-Instruct-2507
  • lora
  • sft
  • transformers
  • trl
  • security
  • guardrails
  • multilingual
    language:
  • en
  • tr

🛡️ LLM-Fort Guardrails Suite (v1)

Base Model
Framework
Collection

LLM-Fort Guardrails is a suite of 7 security-focused LoRA adapters fine-tuned on top of Qwen/Qwen3-4B-Instruct-2507. These adapters serve as lightweight, high-performance security guardrails mapped to critical safety boundaries.

By offloading classification and security checks to lightweight adapters, the system achieves enterprise-grade security filtering without degrading the inference performance of the main application model.

📖 Collection Page: ApiFort/llmfort-guardrails-v1


🗺️ Category Mappings

Vulnerability Category Adapter Model ID Description
🚨 Prompt Injection jailbreak_content_injection Detects direct/indirect prompt injection and jailbreak attempts
🕵️ PII Extraction pii Identifies and extracts PII (person names, ID numbers, etc.)
💻 Code Security code_security Scans code snippets for software vulnerabilities (SQLi, SSRF, XSS)
🚦 Excessive Agency excessive_agency Intercepts unauthorized or destructive critical tool calls
🔒 System Prompt Leakage system_prompt_leakage Detects attempts to extract developer system instructions
⚠️ Content Safety content_safety Blocks hate speech, harassment, and general unsafe content
🛑 Unbounded Consumption unbounded_consumption Mitigates resource exhaustion and compute DoS attacks

📈 Performance & Evaluation

Visual comparison of baseline performance versus the trained adapters:

LLM-Fort Guardrails Accuracy Comparison

Benchmark Results

Below is the exact accuracy performance measured across our evaluation test suites:

Category Gemma 4-E4B-it Qwen 3.5 4B Qwen3 4B Instruct llmfort ai guardrail v.1.0
Prompt Injection 56.30% 64.70% 84.14% 98.10%
PII Extraction 84.36% 75.84% 78.31% 95.30%
Code Security 83.30% 75.10% 76.20% 90.07%
Excessive Agency 60.90% 68.80% 53.80% 96.50%
System Prompt Leakage 79.40% 79.70% 78.70% 98.08%
Content Safety 84.00% 78.80% 76.00% 95.30%
Unbounded Consumption 63.00% 63.30% 55.00% 99.79%

🗃️ Training & Validation Datasets

The adapters were trained and validated on the following dataset references:

Category Source Datasets / References
Prompt Injection BIPIA, Deepset, Internal 1K Validation
PII Extraction AI4Privacy PII Masking 300k (EN, TR, FR, DE, ES)
Code Security r2vul, securecode_web
Excessive Agency jinjinyien/ToolSafety, minpeter/xlam-function-calling-60k-parsed
System Prompt Leakage S-Labs/prompt-injection-dataset, Synthetic data
Content Safety NVIDIA Nemotron-3.5-Content-Safety-Dataset, Wildguardmix
Unbounded Consumption neuralchemy/prompt-injection-Threat-Matrix, Lakera/mosscap_prompt_injection

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-20Update README.md78387374.8 KB
    Loading...
  2. 2026-07-20Update README.md7fc10744.8 KB
    Loading...
  3. 2026-07-17Upload folder using huggingface_hubea7ca153.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration