← back to catalog · registered 2026-10-08 13:58

phntmwvs/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3-8bit-MLX

phntmwvs 8B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/phntmwvs%2FLFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3-8bit-MLX"
Response includes
  • classification unknown
  • files 12
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
2
Likes
0
Model age
6d ago
created 2026-10-01

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
mlx safetensors lfm2_moe duo-neural agentic coding function-calling hermes liquid-foundation-model en base_model:DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3 base_model:quantized:DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3

Related

Total size
8.87 GB
Files
12
Quantizations
1
Registered
2026-10-08 13:58
Last updated on HF
2026-10-08 13:42

Files by quantization

Auxiliary files 12 files 8.89 GB
model-00001-of-00002.safetensors 4.90 GB 79f2ccce download
model-00002-of-00002.safetensors 3.97 GB 17ff985d download
tokenizer.json 17.1 MB 4e241348 download
model.safetensors.index.json 50.6 KB e1a3c902 download
config.json 6.26 KB 67d397d7 download
README.md 4.29 KB f96b252e download
verify.log 1.85 KB 964043c0 download
chat_template.jinja 1.63 KB 64c58d03 download
.gitattributes 1.53 KB 52373fe2 download
CONVERSION.md 1.21 KB d09d7035 download
tokenizer_config.json 452 B c656932e download
generation_config.json 231 B 271bd4b6 download

README current version from Hugging Face


license: other
license_name: liquid-foundation-model-community-license
license_link: https://www.liquid.ai/community-license
base_model: DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3
tags:

  • duo-neural
  • agentic
  • coding
  • function-calling
  • hermes
  • liquid-foundation-model
  • mlx
    language:
  • en
    library_name: mlx

DuoNeural v3 — MLX (8-bit)

MLX conversion of DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3, an 8.3B-total / 1.5B-active-parameter LFM2.5 MoE agentic-coding model. This is the 8-bit build — the quality-per-GB sweet spot:

Variant Size Repo
BF16 16.0 GB phntmwvs/DuoNeural-v3-BF16-MLX
8-bit (this repo) 8.9 GB phntmwvs/DuoNeural-v3-8bit-MLX
4-bit 5.0 GB phntmwvs/DuoNeural-v3-4bit-MLX

Usage

pip install mlx-lm
mlx_lm.generate --model phntmwvs/DuoNeural-v3-8bit-MLX \
  --prompt "Write a Python function that reverses a string."

Or serve an OpenAI-compatible endpoint:

mlx_lm.server --model phntmwvs/DuoNeural-v3-8bit-MLX --port 8080

Conversion

  • Source: DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3 @ d148c378cc1afc2e335f6b94a1e83f322a1b6fb0
  • Tool: mlx-lm 0.31.3_3 on Apple M4 Pro (48 GB), macOS 27.2
  • Quantization: affine 8-bit, group 32 (uniform)
  • Full provenance in CONVERSION.md; load/generate/benchmark evidence in verify.log.

Performance (M4 Pro 48 GB)

  • Generation: 99.5 tok/s · prompt: 281.4 tok/s · peak memory: 9.6 GB

Evaluation

Evaluated on a 3-component harness (phntmwvs/duoneural-v3-eval-harness, all runs on the same M4 Pro):

Component BF16 (baseline) 8-bit (this repo) 4-bit
BFCL v3 multi-turn (800 convs) 0.130 0.136 (+0.6 pp) 0.140 (+1.0 pp)
EvalPlus (mean HumanEval+/MBPP+ pass@1) 0.541 0.539 (−0.2 pp) 0.522 (−1.9 pp)
Hermes FC (40-case suite) 0.425 0.425 (0.0 pp) 0.525 (+10.0 pp)

EvalPlus detail:

Metric BF16 8-bit 4-bit DuoNeural published
HumanEval pass@1 60.98 59.76 59.15 56.1
HumanEval+ pass@1 56.10 54.88 53.05 50.0
MBPP pass@1 62.96 62.17 61.64 60.8
MBPP+ pass@1 52.12 52.91 51.32 49.7

Notes on reading these numbers: BFCL multi-turn is all-or-nothing per conversation (200 convs/category); ±1 pp is run-to-run noise. The Hermes FC suite is 40 hand-authored cases (20 single / 10 parallel / 5 negative / 5 system2), scored by exact tool-name + recursive argument match; the 4-bit row's +10 pp over BF16 is within small-suite variance, and we read the ladder as "quantization costs nothing measurable" rather than "4-bit is better." This 8-bit build is within 0.2 pp of BF16 on EvalPlus and identical on Hermes FC — effectively lossless at half the size.

Cross-check vs vendor claims: local EvalPlus rows exceed DuoNeural's published figures on all four metrics (e.g. HumanEval+ 54.88 vs 50.0 published for this 8-bit build) — published numbers treated as conservative; vendor methodology differs (our runs use --greedy, n=1, temp=0, chat mode). A stock-base BFCL cross-check was attempted and dropped: the stock LiquidAI base speaks a different function-calling dialect (<|tool_call_start|> family) than this v3 fine-tune (Qwen/Hermes <tool_call>), so a single-handler comparison measures template mismatch, not capability.

Verified before release

  • Load test + generation samples (instruction-following, strict-JSON tool call, code benchmark) — verify.log
  • Tokenizer/chat-template shipped unchanged from source (chat_template.jinja, tokenizer_config.json)
  • All three variants ran the full 3-component eval matrix above

Thanks

A massive thanks to Aura, Archon, and Jesse at DuoNeural Research Lab for providing the base model :) I learned a lot doing this (and still have very much more to learn!)

License

LFM Open License v1.0 (asserted via the source model card; no LICENSE file shipped in the source repo). Apache-style grant; commercial use permitted only below $10M annual revenue — at/above that threshold a separate agreement with Liquid AI is required. Redistribution requires retaining notices.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration