license: apache-2.0
base_model:
- huihui-ai/Huihui-Qwen3-14B-abliterated-v2
- Qwen/Qwen3-14B
library_name: openvino
pipeline_tag: text-generation
tags: - openvino
- openvino-genai
- ovms
- intel-npu
- npu
- int4
- qwen3
- qwen3-14b
- abliterated
- text-generation
- local-inference
Huihui-Qwen3-14B-abliterated-v2 OpenVINO INT4 for Intel NPU
This repository contains a ready-to-run OpenVINO IR export ofhuihui-ai/Huihui-Qwen3-14B-abliterated-v2,
prepared for local inference on Intel NPU through OpenVINO Model Server.
It is intended for users who want to skip the local OpenVINO conversion and
weight compression steps.
Source model
- Source model:
huihui-ai/Huihui-Qwen3-14B-abliterated-v2 - Original base model:
Qwen/Qwen3-14B - Architecture:
Qwen3ForCausalLM - Task: text generation
- License: Apache-2.0, inherited from the source model metadata
This is not a fine-tune. It is an OpenVINO INT4 runtime export of the source
model above.
OpenVINO export
The exported directory includes OpenVINO model IR files, OpenVINO tokenizer and
detokenizer files, tokenizer files, chat template, and generation config.
The model config reports max_position_embeddings: 40960.
Tested Intel NPU runtime
Prepared on Windows for OpenVINO / OVMS text generation with:
- Target device:
NPU - OVMS task:
text_generation - Recommended max concurrent sequences:
1 - Recommended cache interval multiplier:
64
Example OVMS command:
ovms.exe `
--model_path Q:/llm/models/OpenVINO/Huihui-Qwen3-14B-abliterated-v2-int4-sym-cw-ov `
--model_name Huihui-Qwen3-14B-abliterated-v2-int4-npu `
--rest_port 8000 `
--rest_bind_address 0.0.0.0 `
--task text_generation `
--target_device NPU `
--max_prompt_len 8192 `
--max_num_seqs 1 `
--cache_interval_multiplier 64 `
--reasoning_parser qwen3 `
--tool_parser hermes3
Local benchmark
This 14B export was uploaded from a local converted OpenVINO INT4 directory, but
it was not separately benchmarked in the upload session. Do not infer token
speed from file size.
Known local artifact size:
openvino_model.bin: about 7.61 GiB
If loading stalls or system memory pressure is too high, reduce--max_prompt_len.