license: apache-2.0
language:
- en
- zh
- multilingual
library_name: transformers
pipeline_tag: text-generation
tags: - 27b
- early-access
- draft
- abliterated
- abliterix
- aeon
- aeon-7
- bf16
- bfloat16
- chat
- coding
- conversational
- function-calling
- gated-deltanet
- gdn
- hybrid-attention
- instruct
- multimodal
- qwen
- qwen3
- qwen3.8
- reasoning
- refusal-removed
- thinking
- tool-calling
- uncensored
- unfiltered
- vision
- vision-language
- vllm
base_model:- AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
base_model_relation: quantized
- AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
Jeethu/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16-PARO
Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
ParoQuant is the state-of-the-art INT4 quantization for LLMs. It closes the accuracy gap with FP16 while running at near-AWQ speed. Supports NVIDIA GPUs (vLLM, Transformers) and Apple Silicon (MLX). For more information, see https://github.com/z-lab/paroquant.
Jeethu/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16-PARO is a 4-bit AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 quantized with ParoQuant.