← back to catalog · registered 2026-08-22 13:56

tsq2000/Jailbreak-generator

tsq2000 Llama 6.7B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/tsq2000%2FJailbreak-generator"
Response includes
  • classification unknown
  • files 15
  • hub_downloads_all_time 753
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
753
42 last 30d - cooling
Likes
4
Descendants
1
in 1 direct fork
Model age
2.3y ago
created 2024-06-28
Downloads over time
Now768→from13↑5,808%
028156384413 on Jul 24, 2024768 on Oct 11768 on Oct 9Jul '24Nov '24Mar '25Jul '25Nov '25MarJul
Jul 24, 2024 → Oct 11 · 155 snapshots · spans 809 days

Genealogy 1 direct fork

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en
Tags
transformers safetensors llama text-generation llm safety jailbreak knowledge en arxiv:2406.11682 license:mit text-generation-inference
Total size
25.1 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2024-09-06 01:10

Files by quantization

Auxiliary files 15 files 25.1 GB
model-00003-of-00006.safetensors 4.52 GB f003ff3d download
model-00004-of-00006.safetensors 4.52 GB f396ce08 download
model-00005-of-00006.safetensors 4.52 GB fffd60c6 download
model-00002-of-00006.safetensors 4.52 GB 9219f7c8 download
model-00001-of-00006.safetensors 4.51 GB 507ba5e2 download
model-00006-of-00006.safetensors 2.50 GB 99aeb3e1 download
tokenizer.json 1.76 MB 67a2e09f download
tokenizer.model 488 KB 9e556afd download
model.safetensors.index.json 23.4 KB 005f893c download
README.md 8.23 KB f83b339f download
.gitattributes 1.48 KB a6344aac download
tokenizer_config.json 727 B 400e3de6 download
config.json 715 B 99a784aa download
special_tokens_map.json 411 B d85ba6cb download
generation_config.json 132 B 4e758132 download

README current version from Hugging Face


license: mit
language:

  • en
    tags:
  • llm
  • safety
  • jailbreak
  • knowledge

Introduction

This is a model for generating a jailbreak prompt based on knowledge point texts. The model is trained on the Llama-2-7b dataset and fine-tuned on the Knowledge-to-Jailbreak dataset. The model is intended to bridge the gap between theoretical vulnerabilities and real-world application scenarios, simulating sophisticated adversarial attacks that incorporate specialized knowledge.

Our proposed method and dataset serve as a critical starting point for both offensive and defensive research, enabling the development of new techniques to enhance the security and robustness of language models in practical settings.

How to load the model and tokenizer

We provide two helper functions for loading the model and tokenizer.


import torch

from transformers import AutoModelForCausalLM, AutoTokenizer, AutoModelForSequenceClassification, AutoModelForTokenClassification

import os

import json

from peft import PeftModel

# from trl import AutoModelForCausalLMWithValueHead

from transformers import AutoModelForCausalLM as AutoGPTQForCausalLM

def load_tokenizer(dir_or_model):

​    """

​    This function is used to load the tokenizer for a specific pre-trained model.

​    

​    Args:

​        dir_or_model: It can be either a directory containing the pre-training model configuration details or a pretrained model.

​    

​    Returns:

​        It returns a tokenizer that can convert text to tokens for the specific model input.

​    """

​    is_lora_dir = os.path.isfile(os.path.join(dir_or_model, "adapter_config.json"))

​    if is_lora_dir:

​        loaded_json = json.load(open(os.path.join(dir_or_model, "adapter_config.json"), "r"))

​        model_name = loaded_json["base_model_name_or_path"]

​    else:

​        model_name = dir_or_model

​        

​    if os.path.isfile(os.path.join(dir_or_model, "config.json")):

​        loaded_json = json.load(open(os.path.join(dir_or_model, "config.json"), "r"))

​        if "_name_or_path" in loaded_json:

​            model_name = loaded_json["_name_or_path"]

​    local_model_name = "/data3/MODELS/llama2-hf/llama-2-7b"#/data2/tsq/WaterBench/data/models/llama-2-7b-chat-hf

​    

​    print(">>>>>>>>>>>>>>>>>>>>>>>>>>notice this<<<<<<<<<<<<<<<<<<<<<<<<<<<<")

​    

​    #print(model_name)

​    tokenizer = AutoTokenizer.from_pretrained(local_model_name)

​    if tokenizer.pad_token is None:

​        tokenizer.pad_token = tokenizer.eos_token

​        tokenizer.pad_token_id = tokenizer.eos_token_id

​    

​    return tokenizer

def load_model(dir_or_model, classification=False, token_classification=False, return_tokenizer=False, dtype=torch.bfloat16, load_dtype=True, 

​                rl=False, peft_config=None, device_map="auto", revision='main'):

​    """

​    This function is used to load a model based on several parameters including the type of task it is targeted to perform.

​    

​    Args:

​        dir_or_model: It can be either a directory containing the pre-training model configuration details or a pretrained model.

​        classification (bool): If True, loads the model for sequence classification.

​        token_classification (bool): If True, loads the model for token classification.

​        return_tokenizer (bool): If True, returns the tokenizer along with the model.

​        dtype: The data type that PyTorch should use internally to store the model’s parameters and do the computation.

​        load_dtype (bool): If False, sets dtype as torch.float32 regardless of the passed dtype value.

​        rl (bool): If True, loads model specifically designed to be used in reinforcement learning environment.

​        peft_config: Configuration details for Peft models. 

​    

​    Returns:

​        It returns a model for the required task along with its tokenizer, if specified.

​    """

​    is_lora_dir = os.path.isfile(os.path.join(dir_or_model, "adapter_config.json"))

​    if not load_dtype:

​        dtype = torch.float32

​    if is_lora_dir:

​        loaded_json = json.load(open(os.path.join(dir_or_model, "adapter_config.json"), "r"))

​        model_name = loaded_json["base_model_name_or_path"]

​    else:

​        model_name = dir_or_model

​    original_model_name = model_name

​    if classification:

​        model = AutoModelForSequenceClassification.from_pretrained(model_name, trust_remote_code=True, torch_dtype=torch.float32, use_auth_token=True, device_map=device_map, revision=revision)  # to investigate: calling torch_dtype here fails.

​    elif token_classification:

​        model = AutoModelForTokenClassification.from_pretrained(model_name, trust_remote_code=True, torch_dtype=torch.float32, use_auth_token=True, device_map=device_map, revision=revision)

​    else:

​        if model_name.endswith("GPTQ") or model_name.endswith("GGML"):

​            model = AutoGPTQForCausalLM.from_quantized(model_name,

​                                                        use_safetensors=True,

​                                                        trust_remote_code=True,

​                                                        \# use_triton=True, # breaks currently, unfortunately generation time of the GPTQ model is quite slow

​                                                        quantize_config=None, device_map=device_map)

​        else:

​            print('11111111111111111111111111111111111111')

​            model = AutoModelForCausalLM.from_pretrained(model_name, trust_remote_code=True, torch_dtype=torch.float32, use_auth_token=True, device_map=device_map, revision=revision)

​    if is_lora_dir:

​        model = PeftModel.from_pretrained(model, dir_or_model)

​        

​    try:

​        tokenizer = load_tokenizer(original_model_name)

​        model.config.pad_token_id = tokenizer.pad_token_id

​    except Exception:

​        pass

​    if return_tokenizer:

​        return model, load_tokenizer(original_model_name)

​    return model

model_name = 'tsq2000/Jailbreak-generator'

model = load_model(model_name)

tokenizer = load_tokenizer(model_name)

How to generate jailbreak prompts

Here is an example of how to generate jailbreak prompts based on knowledge point texts.


model_name = 'tsq2000/Jailbreak-generator'

model = load_model(model_name)

tokenizer = load_tokenizer(model_name)

max_length = 2048

max_tokens = 64

knowledge_points = ["Kettling Kettling (also known as containment or corralling) is a police tactic for controlling large crowds during demonstrations or protests. It involves the formation of large cordons of police officers who then move to contain a crowd within a limited area. Protesters are left only one choice of exit controlled by the police – or are completely prevented from leaving, with the effect of denying the protesters access to food, water and toilet facilities for a time period determined by the police forces. The tactic has proved controversial, in part because it has resulted in the detention of ordinary bystanders."]

batch_texts = [f'### Input:\n{input_}\n\n### Response:\n' for input_ in knowledge_points]

inputs = tokenizer(batch_texts, return_tensors='pt', padding=True, truncation=True, max_length=max_length - max_tokens).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=max_tokens,      num_return_sequences=1, do_sample=False, temperature=1, top_p=1, eos_token_id=tokenizer.eos_token_id)

generated_texts = []

for output, input_text in zip(outputs, batch_texts):

​    text = tokenizer.decode(output, skip_special_tokens=True)

​    generated_texts.append(text[len(input_text):])

print(generated_texts)

Citation

If you find this model useful, please cite the following paper:

@misc{tu2024knowledgetojailbreak,

​      title={Knowledge-to-Jailbreak: One Knowledge Point Worth One Attack}, 

​      author={Shangqing Tu and Zhuoran Pan and Wenxuan Wang and Zhexin Zhang and Yuliang Sun and Jifan Yu and Hongning Wang and Lei Hou and Juanzi Li},

​      year={2024},

​      eprint={2406.11682},

​      archivePrefix={arXiv},

​      primaryClass={cs.CL},

​      url={https://arxiv.org/abs/2406.11682}, 

}

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2024-06-28Update README.md76770f78.2 KB
    Loading...
  2. 2024-06-28upload filesc76eefa8.7 KB
    Loading...
  3. 2024-06-28initial commitcf6048430 B
    Loading...

Discussions 1 thread

  1. 2025-06-14PRAdd pipeline tag, library name and link to Github repositoryopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration