← back to catalog · registered 2026-10-08 13:58

phntmwvs/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3-4bit-MLX

phntmwvs 8B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/phntmwvs%2FLFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3-4bit-MLX"
Response includes
  • classification unknown
  • files 11
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
3
Likes
0
Model age
6d ago
created 2026-10-01

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
mlx safetensors lfm2_moe duo-neural agentic coding function-calling hermes liquid-foundation-model en base_model:DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3 base_model:quantized:DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3

Related

Total size
4.98 GB
Files
11
Quantizations
1
Registered
2026-10-08 13:58
Last updated on HF
2026-10-08 13:42

Files by quantization

Auxiliary files 11 files 4.99 GB
model.safetensors 4.98 GB 5d8bca6e download
tokenizer.json 17.1 MB 4e241348 download
config.json 44.2 KB 01730716 download
model.safetensors.index.json 42.3 KB f465b21f download
README.md 4.44 KB 7730af7c download
verify.log 1.92 KB c4e5da14 download
chat_template.jinja 1.63 KB 64c58d03 download
.gitattributes 1.53 KB 52373fe2 download
CONVERSION.md 1.19 KB f4ecd2ba download
tokenizer_config.json 452 B c656932e download
generation_config.json 231 B 271bd4b6 download

README current version from Hugging Face


license: other
license_name: liquid-foundation-model-community-license
license_link: https://www.liquid.ai/community-license
base_model: DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3
tags:

  • duo-neural
  • agentic
  • coding
  • function-calling
  • hermes
  • liquid-foundation-model
  • mlx
    language:
  • en
    library_name: mlx

DuoNeural v3 — MLX (4-bit)

MLX conversion of DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3, an 8.3B-total / 1.5B-active-parameter LFM2.5 MoE agentic-coding model. This is the 4-bit build — the smallest and fastest, at 5.0 GB it runs comfortably on any 16 GB Apple-silicon Mac:

Variant Size Repo
BF16 16.0 GB phntmwvs/DuoNeural-v3-BF16-MLX
8-bit 8.9 GB phntmwvs/DuoNeural-v3-8bit-MLX
4-bit (this repo) 5.0 GB phntmwvs/DuoNeural-v3-4bit-MLX

Usage

pip install mlx-lm
mlx_lm.generate --model phntmwvs/DuoNeural-v3-4bit-MLX \
  --prompt "Write a Python function that reverses a string."

Or serve an OpenAI-compatible endpoint:

mlx_lm.server --model phntmwvs/DuoNeural-v3-4bit-MLX --port 8080

Conversion

  • Source: DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3 @ d148c378cc1afc2e335f6b94a1e83f322a1b6fb0
  • Tool: mlx-lm 0.31.3_3 on Apple M4 Pro (48 GB), macOS 27.2
  • Quantization: affine 4-bit, group 32; embeddings bumped to 6-bit/g64; MoE router gates held at 8-bit/g64 (per mlx-lm's lfm2_moe quant predicate)
  • Full provenance in CONVERSION.md; load/generate/benchmark evidence in verify.log.

Performance (M4 Pro 48 GB)

  • Generation: 160.9 tok/s · prompt: 146.2 tok/s · peak memory: 5.4 GB

Evaluation

Evaluated on a 3-component harness (phntmwvs/duoneural-v3-eval-harness, all runs on the same M4 Pro):

Component BF16 (baseline) 8-bit 4-bit (this repo)
BFCL v3 multi-turn (800 convs) 0.130 0.136 (+0.6 pp) 0.140 (+1.0 pp)
EvalPlus (mean HumanEval+/MBPP+ pass@1) 0.541 0.539 (−0.2 pp) 0.522 (−1.9 pp)
Hermes FC (40-case suite) 0.425 0.425 (0.0 pp) 0.525 (+10.0 pp)

EvalPlus detail:

Metric BF16 8-bit 4-bit DuoNeural published
HumanEval pass@1 60.98 59.76 59.15 56.1
HumanEval+ pass@1 56.10 54.88 53.05 50.0
MBPP pass@1 62.96 62.17 61.64 60.8
MBPP+ pass@1 52.12 52.91 51.32 49.7

Notes on reading these numbers: BFCL multi-turn is all-or-nothing per conversation (200 convs/category); ±1 pp is run-to-run noise. The Hermes FC suite is 40 hand-authored cases (20 single / 10 parallel / 5 negative / 5 system2), scored by exact tool-name + recursive argument match; this 4-bit row's +10 pp over BF16 is within small-suite variance, and we read the ladder as "quantization costs nothing measurable" rather than "4-bit is better." The real cost shows on EvalPlus (−1.9 pp vs BF16) — still above DuoNeural's published figures on every metric.

Cross-check vs vendor claims: local EvalPlus rows exceed DuoNeural's published figures on all four metrics (e.g. HumanEval+ 53.05 vs 50.0 published for this 4-bit build) — published numbers treated as conservative; vendor methodology differs (our runs use --greedy, n=1, temp=0, chat mode). A stock-base BFCL cross-check was attempted and dropped: the stock LiquidAI base speaks a different function-calling dialect (<|tool_call_start|> family) than this v3 fine-tune (Qwen/Hermes <tool_call>), so a single-handler comparison measures template mismatch, not capability.

Verified before release

  • Load test + generation samples (instruction-following, strict-JSON tool call, code benchmark) — verify.log
  • Tokenizer/chat-template shipped unchanged from source (chat_template.jinja, tokenizer_config.json)
  • All three variants ran the full 3-component eval matrix above

Thanks

A massive thanks to Aura, Archon, and Jesse at DuoNeural Research Lab for providing the base model :) I learned a lot doing this (and still have very much more to learn!)

License

LFM Open License v1.0 (asserted via the source model card; no LICENSE file shipped in the source repo). Apache-style grant; commercial use permitted only below $10M annual revenue — at/above that threshold a separate agreement with Liquid AI is required. Redistribution requires retaining notices.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration