license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.5-9B-abliterated
datasets:
- scottyjmp5/courtlistener-legal-corpus
language: - en
pipeline_tag: image-text-to-text
tags: - legal
- caselaw
- courtlistener
- abliterated
- vision
- tool-calling
- qwen3.5
Legal-Qwen3.5-9B-Abliterated
A legal-domain fine-tune of huihui-ai/Huihui-Qwen3.5-9B-abliterated
(an abliterated Qwen3.5-9B vision-language model), trained on public-domain United States court
opinions from CourtListener. Smaller, faster sibling of
Legal-Qwen3.6-27B-Abliterated.
All base capabilities are preserved and were verified after merging: vision, tool/function
calling, and thinking mode.
Training
- Method: LoRA (r=32, q/k/v/o/gate/up/down) via Unsloth on 1x RTX 5090; merged into the pristine
bf16 base afterward - Data: CourtListener opinions (CPT) + holding-summary pairs (SFT) + 41 recent 2025-2026 rulings
including Chatrie v. United States (S. Ct., June 29, 2026), 2,048-token sequences, 1 epoch
IMPORTANT: pair it with retrieval
Fine-tuning teaches doctrine and style, not verbatim recall. The model can hallucinate reporter
citations, dates, and quotes. For real legal-research use, run it with RAG over the
CourtListener bulk data (or the linked
training corpus) and
verify every citation at the source.
Variants
| Repo | Format | Size | For |
|---|---|---|---|
| this repo | bf16 safetensors | 18 GB | further fine-tuning, vLLM on 24GB+ |
| -GGUF | Q8_0 GGUF + vision projector | 10 GB | Ollama / llama.cpp on 12GB+ |
Speed notes from testing on an RTX 5090 (single-stream): ~47 tok/s bf16 in vLLM with CUDA graphs
off, 87 tok/s with graphs on, 146 tok/s with an FP8 W8A8 quant (llm-compressor FP8_DYNAMIC,
recipe in this card's discussion) plus graphs; ~104 tok/s as Q8_0 in Ollama. Installing the
Triton-based flash-linear-attention package speeds up the hybrid attention layers in vLLM.
Warnings
- Abliterated base: refusal behaviors of the original Qwen3.5 have been removed upstream.
You are responsible for output filtering appropriate to your deployment. - Not legal advice. Outputs are drafts/research aids and must be reviewed by a licensed attorney.
- US-centric: training data is exclusively United States case law.