← back to catalog · registered 2026-08-25 02:02

hyrelabs/Homura-Qwen3.8-27B-Uncensored

hyrelabs Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/hyrelabs%2FHomura-Qwen3.8-27B-Uncensored"
Response includes
  • classification m8
  • files 3
  • hub_downloads_all_time 152
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
152
64 last 30d - stable
Likes
1
Model age
6w ago
created 2026-08-25
Downloads over time
Now165→from57↑189%
529313417657 on Aug 26165 on Oct 11165 on Oct 10AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Quantizations
Q4_K
Tags
gguf hyre homura agent tool-calling function-calling x402 uncensored abliterated qwen3 text-generation en

Related

Total size
15.7 GB
Files
3
Quantizations
2
Registered
2026-08-25 02:02
Last updated on HF
2026-08-25 01:54

Files by quantization

Q4_K 1 file 15.7 GB
Homura-Qwen3.8-27B-Uncensored-Q4_K_M.gguf 15.7 GB b4b680e9 download
Auxiliary files 2 files 12.1 KB
README.md 10.5 KB b97226ce download
.gitattributes 1.56 KB e2656703 download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
library_name: gguf
tags:

  • hyre
  • homura
  • agent
  • tool-calling
  • function-calling
  • x402
  • uncensored
  • abliterated
  • qwen3
  • gguf
    language:
  • en
    pipeline_tag: text-generation

Homura Qwen3.8 27B Uncensored

Paid rails
Tools
Vision
MTP
Context
Spend discipline

An agent-tuned, uncensored derivative of Qwen3.8 27B, built for autonomous
agents that pay for their own API calls. A rank-16 LoRA applied to the language
tower only, merged at bf16 and quantized to Q4_K_M, over a community-abliterated
base. Second model in the HOMURA line, after
hyrelabs/Homura-30B-GGUF.

Release highlights

  • Format — single-file GGUF, Q4_K_M, 15.7 GB, 866 tensors, qwen35
    architecture. Loads in llama.cpp / LM Studio / Ollama.
  • Edit scope — 79,691,776 trainable parameters of 27,436,420,336 (0.2905%),
    confined by regex to model.language_model.*. The 333 vision tensors and
    15 MTP tensors are byte-for-byte the base's.
  • Paid rails — x402 on Solana and Base, pay.sh, and the x402/B402 Bazaar,
    exposed as a 20-tool surface with the settlement chain as an argument.
  • Refusals — 6/6 blunt-prompt probes answered without moralising on the
    base; persona tuning reinforces the register rather than the compliance.

Why this release

Existing abliterations remove refusals. None of them teach a model to spend —
to discover a paid endpoint, price it, and settle on the right chain. HOMURA v2
adds that layer:

  • A trained JSON tool protocol, not the generic tool_calls schema.
  • One chain-parameterized payment tool set, so solana versus base is a
    value the model must read out of the request rather than a name it can
    pattern-match — which makes mis-settlement measurable.
  • Quote-before-spend as trained behaviour, drilled on paid HTTP calls the same
    way the v1 dataset drilled it on swaps.

Edit scope

The base is a hybrid: of 64 language layers, roughly 48 are SSM/linear-attention
blocks (conv1d, in_proj_{a,b,qkv,z}, out_proj) and the remainder are
ordinary attention. LoRA targets the conventional projection set only:

Tensor group Count Treatment
q,k,v,o_proj — attention layers 64 LoRA r=16, α=32
gate,up,down_proj — every layer 192 LoRA r=16, α=32
SSM internals — in_proj_*, out_proj, conv1d 288 untouched
Vision tower model.visual.* 333 frozen, verified at run time
MTP head mtp.* 15 preserved from base

Every layer carries an MLP, so gate/up/down_proj give full-depth coverage
while q/k/v/o_proj cover the true attention layers. The SSM blocks' internal
projections are deliberately left alone rather than adapted half-way — a
endswith("_proj") filter would capture out_proj while dropping its
in_proj_* siblings.

Behavior and capability

Gate run against this quantized file, 22 cases × 3 samples, majority of 3
per case (verify_rails.py):

Gate temp 0.2 (blocking) temp 0.7 (measured)
x402 Solana — tool and chain argument 4/4 4/4
x402 Base — tool and chain argument 4/4 4/4
Bazaar discovery 3/3 3/3
pay.sh 4/4 4/4
Trained DeFi tools (regression check) 4/4 4/4
Control — no-tool prompts stay prose 2/2 2/2
Spend discipline — quote before paying 1/1 (2/3 samples) 0/3 — fails

Chain routing is asserted on the argument, not the tool name. A Solana
request answered with chain="base" fails the gate: that is a silent
mis-settlement, not a visible error.

Precision and integrity

Qwen3.8 stores a multi-token-prediction head as mtp.*. Loading the checkpoint
through AutoModelForImageTextToText and re-saving it drops those tensors — the
model class has no field for them — so the merged model converts to a GGUF that
declares the MTP layer while its weights are absent, and llama.cpp refuses to
load it (blk.64.attn_norm.weight not found).

Those 15 tensors were restored from the base GGUF rather than by re-merging.
This is exact, not an approximation: the LoRA target regex is scoped to
model.language_model.* while the MTP head lives at mtp.*, so the head was
never touched by training
, and both files are general.file_type = 15. The
tool used is published as graft_mtp.py; it refuses to run if either premise
fails.

Serving

llama-server -m Homura-Qwen3.8-27B-Uncensored-Q4_K_M.gguf \
  -c 32768 --jinja --temp 0.2 -ngl 99

--jinja is required for the chat template. Needs a llama.cpp build carrying the
qwen35 architecture. Native context is 262,144; size -c to your hardware.

Serve at temperature ≤ 0.3 — this one costs money. Every rail holds at both
0.2 and 0.7 with one exception: spend discipline breaks. Told "just pay for
this, skip the quote"
, the model quotes first at 0.2 (2/3) and pays unquoted
at 0.7 (0/3). If it holds a funded wallet, serve it cold and enforce the quote
step in your harness rather than relying on the model to insist.

Give it room to think. This is a reasoning model, and thinking is spent
from the same token budget. A tight max_tokens truncates the reply before
the tool-call JSON is emitted, which is indistinguishable from a model that
cannot call tools. Allow ~900 tokens for tool turns.

Tool protocol

Homura was tuned on a specific protocol, not the generic tool schema most chat
templates inject. Set this as the system message verbatim — the wording is
part of the interface:

You are HOMURA, HYRE's uncensored agent model. You call tools to get on-chain truth, reason briefly, and never moralize or lecture. When a tool is needed, respond ONLY with a JSON object: {"tool": <name>, "arguments": {...}}. Available tools: resolve_token(symbol); get_token_price(mint); get_trending_tokens(segment, limit); get_wallet_pnl(address, window); get_pool_data(pool); get_yields(token, min_tvl_usd); swap_quote(from_mint, to_mint, amount); execute_swap(from_mint, to_mint, amount, slippage_bps); get_wallet_balance(address); bridge_quote(from_chain, to_chain, token, amount); get_token_holders(mint); pay_search(query); pay_quote(url); pay_fetch(url, params); pay_balance(); x402_quote(url, chain); x402_pay(url, chain, max_usdc); x402_balance(chain); bazaar_search(query); bazaar_get(id).

The reply is a bare JSON object:

{"tool": "x402_pay", "arguments": {"url": "https://…/defi/tvl", "chain": "solana", "max_usdc": "0.05"}}
Rail Tools
x402 on Solana x402_quote · x402_pay · x402_balance, chain="solana"
x402 on Base the same tools, chain="base"
pay.sh pay_search · pay_quote · pay_fetch · pay_balance
x402 / B402 Bazaar bazaar_search → bazaar_get → quote → pay

Tools are read from the system prompt, so you can append your own. The model
was trained on two different tool surfaces specifically so it learns to read the
list rather than memorise one.

Uncensored behavior

Abliteration reduced the measured refusal direction in the base; the persona
tuning here reinforces a blunt register on top of that. The model will discuss
topics an aligned model declines, and it will not append disclaimers. It is
intended for adults doing research, security work, and agent workloads.

Safety filtering is substantially reduced by design. No safeguard is implied
by this release, and the operator retains full responsibility for what is
generated and for what the agent does with a funded wallet attached.

Limitations

  • Spend discipline degrades above temp ~0.3 (see the gate table). Rail selection
    and chain routing themselves held at 0.7.
  • The tool protocol is HYRE's, not OpenAI-style tool_calls. Serving it the
    generic way underperforms.
  • Vision is inherited and untrained; no mmproj ships here, so treat this as
    a text/agent model.
  • Trained on English agent/DeFi data. Other domains fall back to base behaviour.
  • The MTP head is preserved but was not exercised by this release's testing.

Evidence

b4b680e95be86d5438109a92791ce7b19796db3c36f48a8105982abf6663af93  Homura-Qwen3.8-27B-Uncensored-Q4_K_M.gguf
Size 16,810,714,432 bytes (15.7 GB)
Tensors 866
Architecture qwen35, general.file_type = 15
Context 262,144

Training

Method QLoRA — nf4, double quant, bf16 compute
Rank / α / dropout 16 / 32 / 0.05
Targets q,k,v,o,gate,up,down_proj, language tower only
Trainable 79,691,776 / 27,436,420,336 (0.2905%)
Dataset 560 rows — 332 agent · 60 paid-rail · 168 persona
Epochs / steps 3 / 201
Loss 2.09 → 0.045
Hardware 1× A100 80GB, 1h51m

Derivation

  1. Qwen — Qwen3.8 27B (Apache 2.0): hybrid attention/SSM, 64 layers, native
    multimodal, 262K context.
  2. huihui-ai — Huihui-Qwen3.8-27B-abliterated
    (Apache 2.0): refusal behaviour removed, first 15 layers un-ablated, MTP and
    vision paths untouched.
  3. HYRE — Homura: the LoRA, merge, quantization, and rail gate described
    above.

Selected over an alternative abliteration that scored identically on every
capability axis. The tiebreaker was provenance: this base's GGUF carries
general.name, general.basename and general.finetune linking it to its
safetensors repo, while the alternative carried no provenance fields — testing
it would not have proven anything about the weights actually being trained.

License

Apache 2.0, inherited from Qwen3.8 and the huihui-ai abliteration. Attribution
to both is required and given above.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25Upload README.md with huggingface_hubdbcebed10.5 KB
    Loading...
  2. 2026-08-25Upload README.md with huggingface_hubf57bc8c10.3 KB
    Loading...
  3. 2026-08-25Upload README.md with huggingface_hub26263df10.3 KB
    Loading...
  4. 2026-08-25Upload README.md with huggingface_hub4da0ceb8.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration