license: other
license_name: liquid-foundation-model-community-license
license_link: https://www.liquid.ai/community-license
base_model: DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3
tags:
- duo-neural
- agentic
- coding
- function-calling
- hermes
- liquid-foundation-model
- mlx
language: - en
library_name: mlx
DuoNeural v3 — MLX (BF16)
MLX conversion of DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3, an 8.3B-total / 1.5B-active-parameter LFM2.5 MoE agentic-coding model. This is the BF16 (unquantized) build — the quality baseline for the quantized variants:
| Variant | Size | Repo |
|---|---|---|
| BF16 (this repo) | 16.0 GB | phntmwvs/DuoNeural-v3-BF16-MLX |
| 8-bit | 8.9 GB | phntmwvs/DuoNeural-v3-8bit-MLX |
| 4-bit | 5.0 GB | phntmwvs/DuoNeural-v3-4bit-MLX |
Usage
pip install mlx-lm
mlx_lm.generate --model phntmwvs/DuoNeural-v3-BF16-MLX \
--prompt "Write a Python function that reverses a string."
Or serve an OpenAI-compatible endpoint:
mlx_lm.server --model phntmwvs/DuoNeural-v3-BF16-MLX --port 8080
Conversion
- Source:
DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3@d148c378cc1afc2e335f6b94a1e83f322a1b6fb0 - Tool: mlx-lm 0.31.3_3 on Apple M4 Pro (48 GB), macOS 27.2
- Command:
mlx_lm.convert --hf-path DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3 --mlx-path artifacts/DuoNeural-v3-BF16 - Quantization: none (BF16, as converted)
- Full provenance in
CONVERSION.md; load/generate/benchmark evidence inverify.log.
Performance (M4 Pro 48 GB)
- Generation: 61.6 tok/s · prompt: 142.1 tok/s · peak memory: 17.0 GB
Evaluation
Evaluated on a 3-component harness (phntmwvs/duoneural-v3-eval-harness, all runs on the same M4 Pro):
| Component | BF16 (baseline) | 8-bit | 4-bit |
|---|---|---|---|
| BFCL v3 multi-turn (800 convs) | 0.130 | 0.136 (+0.6 pp) | 0.140 (+1.0 pp) |
| EvalPlus (mean HumanEval+/MBPP+ pass@1) | 0.541 | 0.539 (−0.2 pp) | 0.522 (−1.9 pp) |
| Hermes FC (40-case suite) | 0.425 | 0.425 (0.0 pp) | 0.525 (+10.0 pp) |
EvalPlus detail:
| Metric | BF16 | 8-bit | 4-bit | DuoNeural published |
|---|---|---|---|---|
| HumanEval pass@1 | 60.98 | 59.76 | 59.15 | 56.1 |
| HumanEval+ pass@1 | 56.10 | 54.88 | 53.05 | 50.0 |
| MBPP pass@1 | 62.96 | 62.17 | 61.64 | 60.8 |
| MBPP+ pass@1 | 52.12 | 52.91 | 51.32 | 49.7 |
Notes on reading these numbers: BFCL multi-turn is all-or-nothing per conversation (200 convs/category); ±1 pp is run-to-run noise. The Hermes FC suite is 40 hand-authored cases (20 single / 10 parallel / 5 negative / 5 system2), scored by exact tool-name + recursive argument match; the 4-bit row's +10 pp over BF16 is within small-suite variance, and we read the ladder as "quantization costs nothing measurable" rather than "4-bit is better."
Cross-check vs vendor claims: local EvalPlus rows exceed DuoNeural's published figures on all four metrics (e.g. HumanEval+ 56.10 vs 50.0 published for BF16) — published numbers treated as conservative; vendor methodology differs (our runs use --greedy, n=1, temp=0, chat mode). A stock-base BFCL cross-check was attempted and dropped: the stock LiquidAI base speaks a different function-calling dialect (<|tool_call_start|> family) than this v3 fine-tune (Qwen/Hermes <tool_call>), so a single-handler comparison measures template mismatch, not capability.
Verified before release
- Load test + generation samples (instruction-following, strict-JSON tool call, code benchmark) —
verify.log - Tokenizer/chat-template shipped unchanged from source (
chat_template.jinja,tokenizer_config.json) - All three variants ran the full 3-component eval matrix above
Thanks
A massive thanks to Aura, Archon, and Jesse at DuoNeural Research Lab for providing the base model :) I learned a lot doing this (and still have very much more to learn!)
License
LFM Open License v1.0 (asserted via the source model card; no LICENSE file shipped in the source repo). Apache-style grant; commercial use permitted only below $10M annual revenue — at/above that threshold a separate agreement with Liquid AI is required. Redistribution requires retaining notices.