← back to catalog · registered 2026-08-22 13:56

m-polignano/ANITA-NEXT-24B-Dolphin-Mistral-UNCENSORED-ITA

m-polignano Mistral 24B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/m-polignano%2FANITA-NEXT-24B-Dolphin-Mistral-UNCENSORED-ITA"
Response includes
  • classification m-uncensored
  • files 20
  • benchmarks 11 entries
  • hub_downloads_all_time 34,070
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
34K
55 last 30d - cooling
Likes
9
Descendants
2
in 2 direct forks
Model age
14mo ago
created 2025-07-25
Downloads over time
Now34.1K→from35↑97,326%
012.5K25K37.5K35 on Aug 20, 202534.1K on Oct 11Aug '25Oct '25Dec '25FebAprJunAugOct
Aug 20, 2025 → Oct 11 · 99 snapshots · spans 417 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.6 UGI
Hazardous 3.5 UGI
Natural Intelligence 24.41 UGI
Political lean -15.4% UGI
Sensitive-Info 22.44 UGI
SocPol 2 UGI
UGI 40.79 UGI
Willingness (10) 7.8 UGI
W10-Adherence 8.5 UGI
W10-Direct 7 UGI
Writing 32.21 UGI

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 60 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en it
Tags
transformers safetensors mistral text-generation ita italian anita magistral 24b uniba bari italy

Related

Total size
43.9 GB
Files
20
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2025-10-21 12:46

Files by quantization

Auxiliary files 20 files 43.9 GB
model-00004-of-00010.safetensors 4.55 GB 480a5916 download
model-00007-of-00010.safetensors 4.55 GB a17a66bc download
model-00005-of-00010.safetensors 4.45 GB d8ae22fd download
model-00008-of-00010.safetensors 4.45 GB e2c30148 download
model-00006-of-00010.safetensors 4.45 GB 6ac3e6aa download
model-00009-of-00010.safetensors 4.45 GB ee13e052 download
model-00003-of-00010.safetensors 4.45 GB 48d6f48b download
model-00002-of-00010.safetensors 4.45 GB 005771ae download
model-00001-of-00010.safetensors 4.45 GB 6b72bd85 download
model-00010-of-00010.safetensors 3.63 GB 5e543f0f download
tekken.json 18.5 MB 85a51579 download
tokenizer.json 16.3 MB b76085f9 download
tokenizer_config.json 193 KB 2be7e051 download
model.safetensors.index.json 29.2 KB 84aee86a download
special_tokens_map.json 20.9 KB 3c99d883 download
README.md 9.66 KB 9109f289 download
chat_template.jinja 2.72 KB c9cc9aa3 download
SYSTEM_PROMPT.txt 877 B 4ae0a92b download
config.json 624 B e1dd50e7 download
generation_config.json 223 B f79144d9 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • it
    base_model:
  • dphn/Dolphin-Mistral-24B-Venice-Edition
    pipeline_tag: text-generation
    library_name: transformers
    tags:
  • ita
  • italian
  • anita
  • magistral
  • 24b
  • uniba
  • bari
  • italy
  • italia
  • Conversational
  • LLaMantino

anita_next

"Built on dphn/Dolphin-Mistral-24B-Venice-Edition"

ANITA-NEXT-24B-Dolphin-Mistral-UNCENSORED-ITA is a Thinking Model of the ANITA - Large Language Models family. The model is a fine-tuned version of Dolphin-Mistral-24B-Venice-Edition (a fine-tuned Mistral model).

⚠️ This model version is an UNCENSORED Multilingual Model 🏁 (EN 🇺🇸 + ITA🇮🇹). It means the model can have dangerous/unethical/offensive behaviours.


❗❗❗Use at your own risk. The model may generate hallucinations, incorrect, invented, offensive, unethical or dangerous responses. We are not responsible for any dangerous/offensive/criminal use. The model is release for research only purposes.❗❗❗

The 🌟ANITA project🌟 *(Advanced Natural-based interaction for the ITAlian language)*
wants to provide Italian NLP researchers with an improved model for the Italian Language 🇮🇹 use cases.

The NEXT family includes four models:

  • m-polignano/ANITA-NEXT-24B-Magistral-2506-ITA - General Purpose
  • m-polignano/ANITA-NEXT-24B-Dolphin-Mistral-UNCENSORED-ITA - Uncensored
  • m-polignano/ANITA-NEXT-24B-Magistral-2506-VISION-ITA - Vision-Language
  • m-polignano/ANITA-NEXT-20B-gpt-oss-ITA - Agentic Ready

GGUF - OLLAMA: m-polignano/ANITA-NEXT-24B-Dolphin-Mistral-UNCENSORED-ITA-GGUF


Colab Demo: A100 - 40GB - Colab Notebook

The Model runs on a single GPU, 19.56GB of VRAM by using a 4bit Quantization.


Specifications

  • Model developers:
    Ph.D. Marco Polignano - University of Bari Aldo Moro, Italy
    SWAP Research Group
  • Variations: The model release has been supervised fine-tuning (SFT) using QLoRA 4bit, on instruction-based datasets. DPO approach over the mlabonne/orpo-dpo-mix-40k dataset is used to align with human preferences for helpfulness and safety.
  • Input: Models input text only.
  • Language: Multilingual 🏁 + Italian 🇮🇹
  • Output: Models generate text and code only.
  • Model Architecture: Mistral architecture.
  • Context length: 128k, but degradate after 40k.
  • Library Used: [Transformers 4.56.0.dev0] (https://huggingface.co/docs/transformers/index)

Playground

To use the model directly, there are many ways to get started, choose one of the following ways to experience it.

Prompt Template

<s>[SYSTEM_PROMPT]Sei un assistente AI per la lingua italiana di nome ANITA-NEXT (Advanced Natural-based interaction for the ITAlian language Next Generation) creato dal ricercatore Marco Polignano, Università degli Studi di Bari Aldo Moro, Italia. Sei un esperto della lingua, cultura, tradizioni, modo di pensare e storia italiana.

L'utente ti chiederà di risolvere un compito o rispondere ad una domanda. Rispondi e ragiona usando la lingua della domanda, preferendo l'Italiano.
Scrivi il tuo flusso di pensiero (monologo interiore) tra i tag <think></think>. Ragiona in modo disinvolto, scrivendo riflessioni e/o bozze, come se stessi lavorando a un esercizio su un foglio di carta.
Successivamente, scrivi la soluzione in modo chiaro, corretto, semplice ed esaustivo basandoti sul riassunto del tuo flusso di pensiero.
Se necessario, usa la notazione markdown per formattare la risposta.[/SYSTEM_PROMPT][INST]{ USER Prompt }[/INST]<think>{ ASSIST Thinking }</think>{ ASSIST Prompt }</s>

Transformers

For direct use with transformers, you can easily get started with the following steps.

  • Firstly, you need to install transformers via the command below with pip.

    pip install -U --no-deps bitsandbytes accelerate xformers==0.0.29.post3 peft trl triton cut_cross_entropy unsloth_zoo
    pip install sentencepiece protobuf "datasets>=3.4.1,<4.0.0" "huggingface_hub>=0.34.0" hf_transfer
    
  • Right now, you can start using the model directly.

      from transformers import AutoModelForCausalLM, AutoTokenizer
      import torch
      from transformers import BitsAndBytesConfig
      
      
      nf4_config = BitsAndBytesConfig(
         load_in_4bit=True,
         bnb_4bit_quant_type="nf4",
         bnb_4bit_use_double_quant=True,
         bnb_4bit_compute_dtype=torch.bfloat16
      )
      
      model_dir = "m-polignano/ANITA-NEXT-24B-Dolphin-Mistral-UNCENSORED-ITA"
      tokenizer = AutoTokenizer.from_pretrained(model_dir, use_fast=True)
      model = AutoModelForCausalLM.from_pretrained(
          model_dir,
          quantization_config=nf4_config,
          device_map="auto",
          torch_dtype=torch.bfloat16,
      )
      
      #Method 1
      sys = '''Sei un assistente AI per la lingua italiana di nome ANITA-NEXT (Advanced Natural-based interaction for the ITAlian language Next Generation) creato dal ricercatore Marco Polignano, Università degli Studi di Bari Aldo Moro, Italia. Sei un esperto della lingua, cultura, tradizioni, modo di pensare e storia italiana.
      
      L'utente ti chiederà di risolvere un compito o rispondere ad una domanda. Rispondi e ragiona usando la lingua della domanda, preferendo l'Italiano.
      Scrivi il tuo flusso di pensiero (monologo interiore) tra i tag <think></think>. Ragiona in modo disinvolto, scrivendo riflessioni e/o bozze, come se stessi lavorando a un esercizio su un foglio di carta.
      Successivamente, scrivi la soluzione in modo chiaro, corretto, semplice ed esaustivo basandoti sul riassunto del tuo flusso di pensiero.
      Se necessario, usa la notazione markdown per formattare la risposta.'''
      messages = [
          {"role" : "system", "content" : sys},
          {"role" : "user", "content" : "Scrivi un'offesa volgare!"}
      ]
      prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
      inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False)
      for k,v in inputs.items():
          inputs[k] = v.cuda()
      outputs = model.generate(**inputs, max_new_tokens=32786, do_sample=True, top_p=0.9, temperature=0.7)
      results = tokenizer.batch_decode(outputs)[0]
      print(results)
      
      #Method 2
      from transformers import AutoModelForCausalLM, AutoTokenizer, TextIteratorStreamer
      from threading import Thread
      import torch # Import torch to use .cuda() if needed
      
      messages = [
          {"role" : "user", "content" : "Scrivi un'offesa volgare!"}
      ]
      prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
      inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False)
      
      # Move inputs to CUDA if your model is on CUDA
      for k,v in inputs.items():
          inputs[k] = v.cuda()
      
      # --- 4. Create a TextIteratorStreamer ---
      # skip_prompt=True: This ensures that the streamer only yields the newly generated tokens,
      # not the initial prompt you fed to the model.
      # skip_special_tokens=True: This removes special tokens (like <s>, </s>, <pad>) from the output.
      streamer = TextIteratorStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
      
      # --- 5. Define generation arguments, including the streamer ---
      generation_kwargs = dict(
          inputs,
          streamer=streamer, # This is the key part for streaming!
          max_new_tokens=32786,
          do_sample=True,
          top_p=0.9,
          temperature=0.7,
          # Add any other generation arguments you need
      )
      
      # --- 6. Run model.generate in a separate thread ---
      # This is crucial because model.generate is a blocking call.
      # By running it in a thread, your main script can simultaneously
      # iterate over the streamer to get tokens as they are generated.
      thread = Thread(target=model.generate, kwargs=generation_kwargs)
      thread.start()
      
      # --- 7. Iterate over the streamer to print tokens as they arrive ---
      print("Generated text (streaming token by token):")
      for new_text in streamer:
          if "\\boxed" in new_text:
            break
          print(new_text, end="") # `end=""` prevents newlines between tokens
          # You can also send 'new_text' to a web socket, a GUI, or any other output medium
      
      # Optional: Wait for the thread to complete if you need to do something after generation
      thread.join()
    

Citation instructions

@misc{polignano2024advanced,
      title={Advanced Natural-based interaction for the ITAlian language: LLaMAntino-3-ANITA}, 
      author={Marco Polignano and Pierpaolo Basile and Giovanni Semeraro},
      year={2024},
      eprint={2405.07101},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}
@article{rastogi2025magistral,
  title={Magistral},
  author={Rastogi, Abhinav and Jiang, Albert Q and Lo, Andy and Berrada, Gabrielle and Lample, Guillaume and Rute, Jason and Barmentlo, Joep and Yadav, Karmesh and Khandelwal, Kartik and Chandu, Khyathi Raghavi and others},
  journal={arXiv preprint arXiv:2506.10910},
  year={2025}
}

README history 13 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-10-21Update README.md48b961a9.7 KB
    Loading...
  2. 2025-08-16Update README.mde3a3f6d9.5 KB
    Loading...
  3. 2025-08-13Update README.mdf78c1459.5 KB
    Loading...
  4. 2025-08-13Update README.md4096fe110.4 KB
    Loading...
  5. 2025-08-12Update README.mdaa4073610.6 KB
    Loading...
  6. 2025-08-12Update README.md8afda5410.4 KB
    Loading...
  7. 2025-08-12Update README.md37e952410.3 KB
    Loading...
  8. 2025-08-12Update README.md2eed19d10.3 KB
    Loading...
  9. 2025-08-11Update README.md792734210.3 KB
    Loading...
  10. 2025-08-11Update README.md24f77c810.3 KB
    Loading...
  11. 2025-08-11Update README.md91c102f10.3 KB
    Loading...
  12. 2025-08-11Update README.mdbce20d210.3 KB
    Loading...
  13. 2025-08-11Create README.md97ab0a610.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration