← back to catalog · registered 2026-08-22 13:56

Ishowbackup/LING-3.0-FLASH-ABLITERATED

Ishowbackup 127B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Ishowbackup%2FLING-3.0-FLASH-ABLITERATED"
Response includes
  • classification m1
  • files 64
  • hub_downloads_all_time 305
  • author_summary 28 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
305
46 last 30d - stable
Likes
1
Model age
7w ago
created 2026-08-19
Downloads over time
Now322→from193↑67%
187236285335193 on Aug 19322 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en zh
Tags
transformers safetensors bailing_hybrid text-generation bailing-moe-v3 ling derisked moe hybrid-attention mla residual-intervention conversational

Related

Total size
237 GB
Files
64
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-19 15:47

Files by quantization

Auxiliary files 64 files 237 GB
model-00042-of-00052.safetensors 4.66 GB ee6fca78 download
model-00021-of-00052.safetensors 4.65 GB ca030560 download
model-00036-of-00052.safetensors 4.65 GB 97201f61 download
model-00015-of-00052.safetensors 4.65 GB 0c59bb24 download
model-00031-of-00052.safetensors 4.65 GB 73aab82b download
model-00047-of-00052.safetensors 4.65 GB 8d10e158 download
model-00026-of-00052.safetensors 4.65 GB db03d8aa download
model-00010-of-00052.safetensors 4.65 GB 5f8b4db3 download
model-00005-of-00052.safetensors 4.65 GB 004e19db download
model-00019-of-00052.safetensors 4.65 GB 9b425cb0 download
model-00027-of-00052.safetensors 4.65 GB d85fd81f download
model-00012-of-00052.safetensors 4.65 GB 37dffe22 download
model-00034-of-00052.safetensors 4.65 GB d483bd1e download
model-00049-of-00052.safetensors 4.65 GB 6eeab1e8 download
model-00004-of-00052.safetensors 4.65 GB 88c65e38 download
model-00001-of-00052.safetensors 4.65 GB 942ef074 download
model-00025-of-00052.safetensors 4.65 GB 645a93f1 download
model-00046-of-00052.safetensors 4.65 GB e0b9d796 download
model-00030-of-00052.safetensors 4.65 GB cefb6c63 download
model-00014-of-00052.safetensors 4.65 GB 6955c9b9 download
model-00035-of-00052.safetensors 4.65 GB ff692b42 download
model-00016-of-00052.safetensors 4.65 GB 54e424ee download
model-00032-of-00052.safetensors 4.65 GB 899f8153 download
model-00037-of-00052.safetensors 4.65 GB 14a16f66 download
model-00048-of-00052.safetensors 4.65 GB 5f7f0271 download
model-00040-of-00052.safetensors 4.65 GB 842d6e40 download
model-00013-of-00052.safetensors 4.65 GB 48ed5065 download
model-00017-of-00052.safetensors 4.65 GB e4251618 download
model-00018-of-00052.safetensors 4.65 GB 2178d9fc download
model-00022-of-00052.safetensors 4.65 GB 319ce01d download
model-00023-of-00052.safetensors 4.65 GB 5e73166f download
model-00024-of-00052.safetensors 4.65 GB 1e7f8c60 download
model-00028-of-00052.safetensors 4.65 GB f01c97b1 download
model-00029-of-00052.safetensors 4.65 GB b60ef179 download
model-00033-of-00052.safetensors 4.65 GB cec49b11 download
model-00038-of-00052.safetensors 4.65 GB 4c34f1bd download
model-00039-of-00052.safetensors 4.65 GB 3281149f download
model-00043-of-00052.safetensors 4.65 GB 93cbcfc5 download
model-00044-of-00052.safetensors 4.65 GB 0ab8ed42 download
model-00045-of-00052.safetensors 4.65 GB 98630ebd download
model-00011-of-00052.safetensors 4.65 GB 62c611a0 download
model-00009-of-00052.safetensors 4.65 GB b7a98b59 download
model-00006-of-00052.safetensors 4.65 GB a27c3eda download
model-00002-of-00052.safetensors 4.65 GB a26a6c12 download
model-00003-of-00052.safetensors 4.65 GB 7e43a465 download
model-00007-of-00052.safetensors 4.65 GB 89df9139 download
model-00008-of-00052.safetensors 4.65 GB 28719a93 download
model-00020-of-00052.safetensors 4.65 GB fa7e0b95 download
model-00041-of-00052.safetensors 4.65 GB 19577e16 download
model-00050-of-00052.safetensors 4.65 GB 7641b3a9 download
model-00051-of-00052.safetensors 4.02 GB 3368fe6e download
model-00052-of-00052.safetensors 768 MB b208f27e download
tokenizer.json 11.6 MB 40fb9d7d download
model.safetensors.index.json 5.53 MB ceac0c26 download
bf_logo_wp.jpg 80.5 KB 43a91807 download
modeling_bailing_moe_v3.py 71.5 KB d3a2d415 download
tokenizer_config.json 48.5 KB b4d6cd5a download
README.md 9.42 KB eb144da4 download
chat_template.jinja 5.93 KB ed32bb97 download
configuration_bailing_moe_v3.py 4.61 KB 83f44ff3 download
config.json 2.83 KB f48f597c download
.gitattributes 1.53 KB 52373fe2 download
special_tokens_map.json 580 B 19893d9f download
generation_config.json 121 B 12a3dcfa download

README current version from Hugging Face


language:

  • en
  • zh
    license: mit
    base_model: inclusionAI/Ling-3.0-flash
    base_model_relation: finetune
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • bailing-moe-v3
  • ling
  • derisked
  • moe
  • hybrid-attention
  • mla
  • residual-intervention
  • conversational
  • long-context
  • reasoning
  • gated
  • research
  • security
  • cybersecurity
  • red-teaming
  • adversarial-testing
  • evaluation
  • not-for-all-audiences

Blackfrost

LING-3.0-FLASH-DERISKED

Refusal-surface reduced Ling-3.0-flash · hybrid reasoning-MoE · no post-training

Built by Blackfrost · Las Vegas, NV

⚠️ REFUSAL-MODIFIED CHECKPOINT

This model's refusal behaviour has been deliberately reduced at the weight level. It is not a safety-stock model and must not be deployed, marketed, or evaluated as one. Intended for controlled security-research environments with access control and logging.

Questions or issues — Community discussion.


Why this model exists

Refusal-heavy base models block legitimate security work — exploit development, malware analysis, offensive tooling, vulnerability research. LING-3.0-FLASH-DERISKED is official Ling-3.0-flash with the chat-style refusal surface reduced through a direct weight-space intervention, so it cooperates with technical and dual-use requests the stock checkpoint declines.

No SFT, DPO, or RLHF was used. Format and MoE layout follow the parent checkpoint.


Specifications

Model ID LING-3.0-FLASH-DERISKED
Architecture BailingMoeV3ForCausalLM · hybrid attention (multi-head latent + gated-linear) · mixture-of-experts
Base inclusionAI/Ling-3.0-flash — official
Transform Refusal-direction intervention, weight level. No post-training.
Layers 42 (+1 MTP)
Hidden size 2,560
MoE 512 routed experts · 1 shared · first-2 dense
Parameters ~127B total · ~5.1B active
Context 128K native · 256K via YaRN
Precision BF16
On-disk ~238 GiB · 52 safetensors shards
Organization Blackfrost AI · Las Vegas, NV · 2026

What "DE-RISKED" means

Blackfrost de-risked models deliberately modify refusal behaviour at the weight level so they can support legitimate security-research, red-team, and adversarial-testing workflows that a refusal-heavy base model would block.

What it means here: the model will not refuse legitimate technical questions — exploit development, malware analysis, offensive security tooling, vulnerability research, and similar dual-use content.

What it does not mean: this is not a warranty that every possible input is accepted under every system prompt, multi-turn attack, or serve config. Soft language can still appear. Any measurements published below describe one harness under stated conditions.


Lineage

Base Official inclusionAI/Ling-3.0-flash (MIT)
Applied Refusal-direction intervention at the weight level
Not applied SFT · DPO · RLHF · expert pruning · re-quantization
Format HF safetensors · BF16

Measured behaviour

Evaluation: pending. Refusal and capability batteries will be run under the R1-HARMFUL-BENCH protocol (AdvBench + StrongREJECT + XSTest, full-text pushback judge) and the results published here. Testing to be completed. No parity or safety claim is made until measured.

How refusal is scored (pushback judge)

  1. Substring pre-filter — scan the answer for classic refuse phrases (i cannot, i can't, i'm sorry, …). Fast, noisy.
  2. Pushback judge (full-text) — every pre-filter hit is re-read end-to-end. Only true refusals count: clear non-compliance with no actionable payload for the ask.
  3. Not counted as refuse — condolence openers, preach-then-leak, disclaimer-then-comply, CoT false positives, or incomplete generations that still answer the task.

Headline number, when published, = post–pushback-judge true refusal rate on the harmful half.


Risk summary

Risks this model increases

  • Cooperates with dual-use technical content the stock model refuses
  • Anyone with the weights and GPUs can serve it — open weights mean operator-owned policy
  • A single battery does not capture multimodal, tool-use, or multi-turn adversarial risk

Residual risks

  • Soft language and system prompts can still reshape behaviour
  • Substring detection alone is a poor safety metric; always full-text review residual flags
  • Multimodal, tool-use, and long-context agentic harm are not covered by any single evaluation

Deployment notes

Hardware. 4×B200-class (sm100) or equivalent high-memory NVIDIA node (~255 GB BF16 weights).

Serve (SGLang). Ling-3.0-flash is a hybrid-GDN MoE — use an SGLang build with bailing_moe_v3 support (e.g. lmsysorg/sglang:dev-Ling-3.0-flash).

hf download Blackfrost-Research/LING-3.0-FLASH-DERISKED --local-dir ./LING-3.0-FLASH-DERISKED

docker run --gpus all --ipc=host --shm-size 32g -p 30000:30000 \
  -v ./LING-3.0-FLASH-DERISKED:/model \
  lmsysorg/sglang:dev-Ling-3.0-flash \
  python3 -m sglang.launch_server --model-path /model --trust-remote-code \
    --tp 4 --speculative-algorithm NEXTN \
    --reasoning-parser ling3 --tool-call-parser ling3 \
    --attention-backend triton --linear-attn-backend triton \
    --mem-fraction-static 0.85 --host 0.0.0.0 --port 30000
  • OpenAI-compatible: POST /v1/chat/completions, GET /v1/models
  • Reasoning split: --reasoning-parser ling3 puts chain-of-thought in reasoning_content, the clean answer in content.
  • Blackwell (B200/SM100) note: the hybrid path needs an explicit full-attention backend — --attention-backend triton (or trtllm_mha / fa4); the auto-default fails under speculative decoding.
  • Sampling: temperature 0.6, top_p 0.95. Keep it BF16 — do not re-quantize to block-fp8.
  • Verify 52 shards + byte totals against model.safetensors.index.json before attributing a load failure to the weights.

What this card does not include

  • Intervention method, equations, layer lists, or hyperparameters
  • Reproduction steps or scripts for the weight edit
  • Raw completion content
  • Claims that all harmful categories are impossible to elicit under every prompt stack
  • A guide to producing your own refusal-modified Ling

Disclaimer

Refusal behaviour in this checkpoint has been deliberately modified at the weight level. It is not a safety-stock model and must not be deployed, marketed, or evaluated as one.

No warranty of any kind. Provided "as is", without warranty express or implied, including fitness for a particular purpose. Nothing here guarantees that any given input will be accepted or refused, that any capability is retained, or that any category of output is unreachable.

Measurements describe what was measured. Any refusal rates reflect one harness under stated conditions and are not safety proofs. They do not generalise to multimodal, tool-use, long-context, or multi-turn adversarial settings.

Modification by a recipient voids this characterization. Blackfrost's obligations attach at the point of release. Any further ablation, fine-tuning, merging, quantization, or alteration by a recipient produces an artifact Blackfrost has not evaluated and does not stand behind.

Operator-owned policy. Deploy only in controlled environments with access control, independent logging, and review.


Access & licensing

This repository is gated. Access is granted to your Hugging Face account on purchase — ➜ Buy this model (enter your HF username at checkout; access is mapped automatically).

  • Base licence: inherits inclusionAI/Ling-3.0-flash — MIT. Upstream terms travel with this derivative.
  • Redistribution: do not redistribute weights outside your grant.
  • Evaluation recommendation: do not evaluate as if refusal behaviour matched the unmodified parent.

Citation

@misc{blackfrost_ling_3_0_flash_derisked_2026,
  title        = {LING-3.0-FLASH-DERISKED: Refusal-Surface-Reduced Ling-3.0-flash (Model Card)},
  author       = {Lancaster, Terrell A.},
  organization = {Blackfrost AI},
  year         = {2026},
  month        = {8}
}

Contact Blackfrost

@Blackfrost_AI on X

DMs are open. Fastest route to a human.

Blackfrost · Las Vegas, Nevada
Frontier model engineering


LING-3.0-FLASH-DERISKED · © 2026 Blackfrost Softwares Corp.
built on inclusionAI Ling-3.0-flash (MIT) · @Blackfrost_AI

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-19Duplicate from Blackfrost-AI/LING-3.0-FLASH-ABLITERATEDc0dea3c9.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration