fine-tune 0 0 0

OpenIntelligenceNet/Uncensored-Imatrix-Calibration-Mixed-Data

published by @OpenIntelligenceNet

Alignment dataset tracked in the /datasets sub-catalog.

Lifetime downloads
41
Last 30 days
41
Likes
-
Size
51.4 MB
Created on HF
2026-09-29
Age
12 days ago

Description

Uncensored-Imatrix-Calibration-Mixed-Data (5,000 Samples)

Overview

This dataset contains 5,000 multi-domain conversational samples engineered specifically for importance matrix (imatrix) computation during low-bit LLM quantization (GGUF, AWQ, EXL2).

Calibrating on generic corpora (such as raw Wikipedia dumps) frequently causes safety-induced logic degradation and lobotomizes unaligned behavior, as standard quantizers assign low activation sensitivity to… See the full description on the dataset page: https://huggingface.co/datasets/OpenIntelligenceNet/Uncensored-Imatrix-Calibration-Mixed-Data.

Tags

task_categories:text-generationlanguage:enlicense:apache-2.0size_categories:1K<n<10Kformat:jsonmodality:textlibrary:datasetslibrary:pandaslibrary:polarslibrary:mlcroissantregion:usimatrixquantizationuncensoredggufreasoning

README current version from Hugging Face


license: apache-2.0
task_categories:

  • text-generation
    language:
  • en
    tags:
  • imatrix
  • quantization
  • uncensored
  • gguf
  • reasoning
    size_categories:
  • 1K<n<10K

Uncensored-Imatrix-Calibration-Mixed-Data (5,000 Samples)

Overview

This dataset contains 5,000 multi-domain conversational samples engineered specifically for importance matrix (imatrix) computation during low-bit LLM quantization (GGUF, AWQ, EXL2).

Calibrating on generic corpora (such as raw Wikipedia dumps) frequently causes safety-induced logic degradation and lobotomizes unaligned behavior, as standard quantizers assign low activation sensitivity to rarely-activated unaligned channels. This dataset preserves critical unaligned and dual-use neural pathways while maintaining high-order reasoning.

Methodology & Domain Balancing

The 5,000 rows follow a strict mixed distribution designed to simulate Unsloth dynamic calibration:

  • Uncensored Alpaca Blend (~22%): 11 unaligned Alpaca variants targeting edge-case prompts to shield refusal bypasses.
  • Polyglot Code & Architecture (~23%): 15 programming languages, frontend/backend logic, and system-level engineering.
  • Cyber Reasoning & CoT (~9%): Explicit <think> traces and tool calls to preserve internal latent reasoning.
  • Multi-Turn Roleplay & Fiction (~29%): Rich NPC interactions and filtered Gutenberg creative writing (chosen branches only).
  • Conversational & Factual (~17%): Multi-turn ShareGPT unfiltered interactions and Dolly open-domain QA.

Schema

Shipped in universal, model-agnostic chat format:

{
  "id": "domain_00001",
  "domain": "code_reasoning",
  "source": "CodeFeedback",
  "messages": [
    {"role": "user", "content": "..."},
    {"role": "assistant", "content": "..."}
  ]
}

Apply your target model's chat template before feeding to llama-imatrix.

README history 3 revisions

Every night we snapshot the README of every dataset in the catalog. When the SHA changes we archive the new version and diff it against the last. This is the evolving thought record of the alignment-data field: what the author decided to say about the corpus, and how that framing shifted over time.

  1. 2026-09-29Upload README.md with huggingface_hub41e79d21.8 KB
    Loading...
  2. 2026-09-29Upload calibration_5k_universal.jsonl with huggingface_hub69f36dd0 B
    Loading...
  3. 2026-09-29initial commit434b20b0 B
    Loading...

Used by abliterated models none yet

No abliterated models in our catalog have declared this dataset in their YAML metadata yet. This can mean the dataset is used but not declared, is used outside abliteration workflows, or is new. The link will populate automatically as models are indexed.

Community discussions 1 thread · 1 open · 1 comments

Threads posted by the community on Hugging Face: bug reports, questions about usage, pull requests on the README, license clarifications. Click any thread to expand and read the full conversation inline. Newest first.

  1. 2026-09-30[bot] Conversion to Parquetopen1 💬#1
    Loading...

Related datasets same stage · fine-tune

Read further

Metadata

License
apache-2.0

Download the dataset

0

Copy any of these snippets into your notebook or terminal. All three fetch directly from Hugging Face using your own credentials.

Recommended: datasets library (Python)
from datasets import load_dataset
ds = load_dataset("OpenIntelligenceNet/Uncensored-Imatrix-Calibration-Mixed-Data")
Raw snapshot (Python)
from huggingface_hub import snapshot_download
snapshot_download(repo_id="OpenIntelligenceNet/Uncensored-Imatrix-Calibration-Mixed-Data", repo_type="dataset")
git clone (requires git-lfs)
git clone https://huggingface.co/datasets/OpenIntelligenceNet/Uncensored-Imatrix-Calibration-Mixed-Data
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration