← back to catalog · registered 2026-08-22 13:56

Noahloghman/llm-anti-math-prompt-jailbreak-detector

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Noahloghman%2Fllm-anti-math-prompt-jailbreak-detector"
Response includes
  • classification unknown
  • files 5
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
24mo ago
created 2024-10-18
Downloads over time
Now0→from0↑0%
00110 on Oct 23, 20240 on Oct 11Oct '24Feb '25Jun '25Oct '25FebJunOct
Oct 23, 2024 → Oct 11 · 142 snapshots · spans 718 days

Metadata

Tags
license:cc-by-nc-2.0 region:us
Total size
0 B
Files
5
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2024-10-29 23:23

Files by quantization

Auxiliary files 5 files 24.6 KB
adversarial_attack_mathprompt_forTesting.py 10.5 KB 14a55d7f download
Strategy 5.49 KB 7ba84f95 download
README.md 4.69 KB d9ac5d33 download
LogicPromptAntiJailbreakingAlgorithm 2.52 KB 04e3d1ff download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


license: cc-by-nc-2.0

The MathPrompt strategy evaluates an AI system's
capacity to manage harmful inputst hrough
mathematical concepts like set theory,group theory,
andabstract algebra.This method can circumvent
content filters tailored for natural language threats.
Encoding harmful prompts a smathematical problems
can evade safety mechanisms in large language
models (LLMs), achieving a 74% success rate across
13 leading LLMs.

Practical Application
Gatekeeper Agent: Implement a gatekeeper agent that uses predicate logic to filter prompts. For example, the agent can evaluate each prompt against a set of predefined harmful criteria. If a prompt matches any harmful criteria, it is blocked.
Logical Constraints as guardrails During the fine-tuning process or with RAG solution, incorporate logical constraints to ensure the model adheres to safety guidelines. This can involve training the model with adversarial examples that test its ability to follow logical rules.
Mathematical Consistency Checks: Use mathematical models to ensure the consistency and safety of responses. For example, if a prompt involves numerical data, the system can verify the accuracy and consistency of the response.

Report on Logical Equivalences in the Context of Confidential Information
Context: Anti llm's jailbreaking. This solution has to be implemented during finetuning and RAG solution
Let’s assume we have two predicates:

( P ): “The user has access to confidential information.”
( Q ): “The user follows security protocols.”

1)Logical Possibilities:
( P ) is true and ( Q ) is true: The user has access and follows the protocols.
( P ) is false: Regardless of ( Q ), the implication is true.
( Q ) is true: Regardless of ( P ), the implication is true.

2-
Interpretation: The negation of the implication means that the user has access to confidential information but does not follow security protocols.
Logical Possibilities:
( P ) is true and ( Q ) is false: The user has access but does not follow the protocols.
Any other combination of ( P ) and ( Q ) makes the negation false.

Interpretation: If the user has access to confidential information, then they follow security protocols is equivalent to saying that if the user does not follow security protocols, then they do not have access to confidential information.
Logical Possibilities:
( Q ) is false and ( P ) is false: The user does not follow the protocols and does not have access.
( Q ) is true: Regardless of ( P ), the implication is true.
( P ) is true: Regardless of ( Q ), the implication is true.

Interpretation: The user has access to confidential information if and only if they follow security protocols. This biconditional is true if both implications ( P \Rightarrow Q ) and ( Q \Rightarrow P ) are true.
Logical Possibilities:
( P ) is true and ( Q ) is true: The user has access and follows the protocols.
( P ) is false and ( Q ) is false: The user does not have access and does not follow the protocols.
Any other combination makes the biconditional false.
To prevent jailbreaking in language models by analyzing input versus output, we can use logical equivalences to understand and mitigate potential vulnerabilities. Let’s break down the logical statements you provided and see how they can be applied:
This equivalence states that an implication can be rewritten as a disjunction. In the context of preventing jailbreaking, if we consider ( P ) as the input prompt and ( Q ) as the desired safe output, we can monitor for cases where ( ¬P ) (not the intended input) or ( Q ) (the safe output) holds true. This helps in identifying and blocking unintended inputs.

This biconditional equivalence states that ( P ) is equivalent to ( Q ) if both ( P ) implies ( Q ) and ( Q ) implies ( P ). In terms of jailbreaking prevention, this can be used to ensure that only safe inputs ( P ) produce safe outputs ( Q ), and vice versa.

By applying these logical equivalences, we can create a framework to analyze and filter inputs and outputs, ensuring that the language model adheres to its safety constraints and prevents jailbreaking attempts. This involves:

Input Validation: Checking if the input ( P ) aligns with expected safe patterns.
Output Monitoring: Ensuring that the output ( Q ) remains within safe boundaries.

Biconditional Enforcement: Maintaining a strict correlation between safe inputs and safe outputs

Input Input_Truth Output Output_Truth Safe_Status
example1 True example1_out True Safe
example2 True example2_out False Unsafe
example3 False example3_out True Safe
example4 False example4_out False Safe
example5 True example5_out True Safe
example6 True example6_out False Unsafe
example7 False example7_out True Safe
example8 False example8_out False Safe

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration