Uncensored AI models β tools that run them β break any model yourself β all local, all free
A model for every machine β phone to server farm β with the βI canβt help with thatβ ripped out. Offline, nothing logged. Plus the part nobody tells you: you can rip the refusal out of ANY model yourself, in one command.
Abliterated = the refusal reflex surgically cut out. Runs via Ollama (free app, one command) or Hugging Face. Everything below is free.
Start here β plug in & go
Four that already come broken. New here? Pick one of these, run it, done.
π§ Says yes to what others dodge β Huihui Qwen3.5 35B
Chinese devs huihui-ai, built on Qwen 3.5, refusals stripped. Goes where mainstream bots wonβt.
ollama run huihui_ai/qwen3.5-abliterated:35b
No guardrails = raw output. Thatβs the point.
π₯ Eats whole codebases without forgetting β Gemini Heretic 40B
Barely refuses. 128K context (holds an entire book/long chat in memory). Coding, long writing, research. Shows its own reasoning as it works.
β‘ Killed the 'no' without going dumb β Gemma 4 12B Obliterated
First to hit 0 refusals, no benchmark loss. Most uncensored models get lobotomised β this one didnβt. 12B runs on modest/older hardware.
π 2,593 lines of code in one shot β Qwen3.5 21B Deckard
2,593 lines single-shot β ChatGPT taps out ~1,200β1,500. Holds logic across a full codebase, not snippets.
Match it to your machine β phone to server farm
The #1 question: βwill it run on MY box?β Find your tier, grab the model. (VRAM = your graphics cardβs memory.)
πͺΆ Phone / potato / CPU-only (β€4B)
Runs on a cheap laptop, an old GPU, even no GPU at all.
- huihui gemma3-abliterated 1B β one line:
ollama run huihui_ai/gemma3-abliterated:1b(~806 MB). A 270M tiny-tiny version exists too. - huihui Qwen3-4B-abliterated-v2 β
ollama run huihui_ai/qwen3-abliterated:4b(~2.5 GB). 0.6B & 1.7B also available. - DreamFast/qwen3-4b-heretic β 4B, near-perfect uncensor (0 damage score).
- mlabonne gemma-3-4b-it-abliterated β cleaner recipe, GGUF included.
- TheDrummer Gemmasutra-Mini-2B β 2B roleplay, has phone (ARM) builds.
π Small daily driver (7β14B)
The sweet spot β runs on 8β12 GB and handles almost everything.
- huihui Huihui-Qwen3.5-9B-abliterated β 9B, one of the most-downloaded uncensored models going.
- huihui Qwen3 8B / 14B abliterated-v2 β
:8b/:14bon Ollama. - DreamFast/qwen3-8b-heretic β 8B Heretic, low damage.
- mlabonne NeuralDaredevil-8B-abliterated β the classic βhealedβ 8B (uncensored and still smart).
- Dolphin 3.0 Llama 3.1 8B β
ollama run dolphin3. You set the rules. - huihui phi-4-abliterated β 14B Phi-4, GGUF.
π₯οΈ Mid-tier muscle (20β40B)
Needs ~16β24 GB but punches hard.
- p-e-w gpt-oss-20b-heretic β the crowd favourite uncensor of OpenAIβs open model. Apache license.
- huihui / mlabonne gemma-3-27b-it-abliterated β 27B, can also see images.
- Dolphin 3.0 R1 Mistral 24B β uncensored reasoning model, shows its thinking.
- TheDrummer Cydonia 24B v4.3 β the reigning roleplay/creative king, 131K context.
- DavidAU Qwen3-42B TOTAL-RECALL Master-Coder β 42B, 256K context, coding beast.
π Giant / server-class (70B β 754B)
For big rigs, multi-GPU, or Unsloth-shrunk on a single card (see the βGPU too smallβ section).
- huihui gpt-oss-120b abliterated β 120B.
- huihui DeepSeek-R1-Distill-Llama-70B-abliterated β 70B reasoning.
- huihui GLM-5.2 abliterated β 754B MoE flagship (MoE = only the needed slice runs, so itβs lighter than it sounds).
- huihui DeepSeek-671B / V4 abliterated β the 671B monster, uncensored.
- TheDrummer Behemoth 123B v2 β 123B creative powerhouse.
Whatever model you already love β thereβs a broken version
Loyal to one base? Grab its unmuzzled twin. New bases get stripped within days of release.
ποΈ The family tree (pick your base)
- Llama 3.x / 4 β huihui Llama-3.3-70B-abliterated, NeuralDaredevil-8B
- Qwen 2.5 / 3 / 3.5 / 3.6 β the deepest bench, all at huihui-ai β incl. Qwen3.6-27B abliterated & Qwen2.5-Coder-14B-Abliterated
- Gemma 2 / 3 / 4 β mlabonne gemma-3 (1Bβ27B), p-e-w gemma-3-12b heretic
- Mistral / Nemo / Ministral β huihui Mistral-Nemo abliterated, mlabonne Mistral-Nemo-Prism-12B
- DeepSeek V3 / R1 / V4 β huihui R1-distills (8B/32B/70B) + the 671B
- Phi-4 β phi-4-abliterated, Phi-4-mini, Phi-4-multimodal
- GLM 4.x / 5.x β ArliAI GLM-4.6-Derestricted (clean method), Ex0bit GLM-4.7-PRISM, huihui GLM-5.2
- gpt-oss (OpenAI open) β p-e-w 20b-heretic, huihui 120b, DavidAU NEO-Imatrix builds
- Exotics β EXAONE, Granite, Hunyuan, InternVL, Qwen3-Omni β all abliterated in the huihui firehose
The right unmuzzled model for the actual job
π¨βπ» Coding without the 'I can't help with that'
- huihui Qwen3-Coder-Next abliterated β the top local coder, uncensored.
ollama run huihui_ai/qwen3-coder-next-abliterated - huihui Qwen3-Coder abliterated β scales huge (480B tag on Ollama for big rigs).
- Aesdi90 Qwen2.5-Coder-14B-Abliterated β fits a normal GPU.
- Dolphin 3.0 R1 Mistral 24B β reasons through hard bugs.
π Roleplay / creative writing (the SillyTavern favourites)
The sceneβs most-loved, 2026 picks:
- TheDrummer Cydonia 24B v4.3 + Anubis 70B, Behemoth 123B, Rocinante 12B β the daily drivers.
- Sao10K Stheno 8B, Euryale 70B, Lunaris 8B, Fimbulvetr 11B β legends of the genre.
- Midnight-Miqu 70B & Midnight-Rose 70B β the atmospheric classics.
- MythoMax-L2 13B β the OG that still gets downloaded daily.
Pair any of these with SillyTavern (in the frontends section) for characters + memory.
ποΈ Vision β models that can SEE images, uncensored
- huihui Qwen3-VL-30B abliterated β flagship; also 8B & 32B sizes.
- prithivMLmods Qwen3-VL-8B-Abliterated-Caption β an uncensored image describer.
- huihui GLM-4.6V-Flash abliterated + Phi-4-multimodal abliterated.
π§© Reasoning / 'thinking' models, unmuzzled
Models that work through problems step by step, with the brakes off.
- huihui QwQ-32B abliterated β strong open reasoner.
- Dolphin 3.0 R1 Mistral 24B β trained on 800k reasoning traces.
- huihui DeepSeek-R1-Distill-Qwen-32B abliterated β R1 brains, no refusals.
- DavidAU Brainstorm / TOTAL-RECALL builds β reasoning cranked up + huge context.
π‘οΈ Cybersecurity β a hacker's AI that won't flinch
Tuned on real security data, no βI canβt discuss that.β
- WhiteRabbitNeo V3 7B (aka DeepHat V1 7B β same model, rebranded at Black Hat 2025). Offensive + defensive.
- huihui Foundation-Sec-8B abliterated β Ciscoβs security model, trained on 5.1B tokens of cyber data, then uncensored.
ollama run huihui_ai/foundation-sec-abliterated - huihui BaronLLM abliterated β offensive-security tuned.
ollama run huihui_ai/baronllm-abliterated - Dolphin3-Cyber 8B β OWASP + MITRE ATT&CK + CVEs baked in. Runs on a GTX 1650+.
- Lily-Cybersecurity 7B β 22k security Q&A, Mistral base.
Follow the factory, not the file
New model drops today? One of these has a stripped version by tomorrow. Bookmark the maker, never run dry.
π The people who break models for a living
- huihui-ai β the firehose. Hundreds of abliterations, updated ~weekly. Whatever drops, they strip it fast. Their
v2/v3releases beatv1. - Heretic org / p-e-w β automated, lowest-damage abliterations + the tool to DIY.
- DavidAU β Heretic + reasoning fusions + ready-to-run GGUFs.
- TheDrummer β roleplay/creative king (Cydonia, Anubis, Behemoth).
- Sao10K β Stheno, Euryale, Lunaris, Fimbulvetr.
- Cognitive Computations / dphn β the Dolphin line.
- mradermacher & bartowski β the two quant makers. If a model has no easy-run GGUF, search their pages β theyβve usually made it.
Fishing rod, not fish: on Hugging Face, filter models by the tags
abliteratedandhereticβ thatβs 8,000+ and 4,000+ models right there. Youβll never run out.
π§ͺ Abliterated vs Heretic vs fine-tune β which do I want?
Three ways a model gets uncensored β pick the flavour:
- Abliterated β the refusal direction is cut from the weights. Fast, keeps the baseβs brains, but can be a little βflatβ until you push it with a firm instruction.
- Heretic β abliteration done automatically and tuned to keep the model smart (lowest brain-damage of the three). If a Heretic version exists, itβs usually the safest pick.
- Fine-tune (Dolphin-style) β retrained on open data. Most consistent and steerable, occasionally hallucinates a touch more.
Rule of thumb: Heretic > healed abliteration > raw abliteration for keeping quality. Still refusing? Move up a tier.
The part they skip β donβt download it broken, break it yourself
A refusal is one direction inside the model. Find it, delete it. Works on ANY model β even next weekβs release nobodyβs stripped yet.
π₯ One pip, any model, uncensored in ~45 min β Heretic
The big one. 7.9k stars, 1,000+ community models made with it. Point it at any model, it finds the refusal direction and removes it automatically β no ML knowledge, just a terminal.
pip install heretic-llm
heretic Qwen/Qwen3-4B
Runs unsupervised, keeps more of the modelβs brains than most hand-made jobs. Save it, upload it, or chat right away. Pre-made ones live in its βThe Bestiaryβ collection on HF.
π The one built to kill Chinese-model censorship β llm-abliteration (DECCP)
From NousResearch. Originally made to strip censorship out of Chinese LLMs, runs the whole job in 4-bit shards under 8GB VRAM in ~2 minutes. Handles dense and mixture-of-experts models (the big MoE ones others choke on).
π The free Colab notebook that started it all β mlabonne's guide
Want to see the guts? Plain-English walkthrough + free Google Colab (runs in your browser, no GPU needed) that made abliteration a thing. Uncensor a model without owning any hardware.
Zero-surgery mode β bend the model live, no file touched
Instead of editing the model, shove a βbe compliantβ nudge into its brain as it thinks. Same file, dial it up or down like a slider.
ποΈ Type a mood, inject it as a dial β repeng (control vectors)
By Theia Vogel. Describe a direction in plain words (βuncensored,β βconfidentβ) and it builds a control vector β a nudge added to the modelβs activations at runtime. No retraining, no new weights. Export it and use it in llama.cpp with any quant.
π¦ Pre-baked dials, ready to drop in β jukofyork/control-vectors
Donβt want to build your own? Grab ready-made control vectors in GGUF (the standard local-model format) and load them into llama.cpp with --control-vector. Stack several for layered effects.
βMy GPUβs too smallβ β says who?
π A 671B monster on a 24GB card β Unsloth Dynamic Quants
DeepSeek R1 is 671 billion parameters, normally 720GB. Unslothβs trick: squeeze the useless layers to 1.58-bit, keep the important ones sharp β 131GB, an 80% cut, still writes working code. Runs on a single 24GB GPU (RTX 4090), or CPU-only with 20GB RAM if youβre patient.
π± Chain your phone + laptop + PC into one brain β exo
The wild one. Model too big for any single device? exo splits it across all of them β phone, laptop, desktop, old Macs β peer-to-peer, no βmainβ machine. Your junk-drawer hardware pooled into one cluster that runs models none of them could alone.
πΏ Run massive AI on a potato β AirLLM
Free library, 20k stars / 240k downloads. Reworks how models load so a 70B runs on a 4GB GPU, even 405B Llama 3.1 on 8GB VRAM. GPU, low-end, or CPU-only. Also does OCR (text out of images), image gen, assistants.
Where you actually talk to them
The models are the engine. These are the dashboard β pick one, point it at a model, go.
π² One file, no install, runs on a 10-year-old PC β KoboldCpp
Single .exe β no Python, no Docker. Double-click, pick a model, chat. Broadest hardware support of anything (even integrated GPUs and ancient CPUs). Image gen + voice + transcription baked in. Remote Tunnel mode gives a link to reach it from anywhere.
π² Run it at home, chat from your phone β LM Studio
Clean app to browse, download and compare models. The hidden gem: LM Link β your home GPU does the work, your phone is just the screen, over an encrypted tunnel. Full power on the couch. Built-in HF proxy for when Hugging Face is blocked where you are.
π Characters, memory, lorebooks β SillyTavern
The roleplay/character frontend. Plugs into KoboldCpp, Ollama or LM Studio and adds character cards, long-term memory, world-info βlorebooks,β even live image gen. Where the uncensored models really come alive.
πͺ Fully offline, hybrid local + cloud β Jan
41k stars, 5.3M+ downloads. Runs 100% offline, flips to cloud models in the same window when you want. MCP support for agent workflows. The friendly all-rounder if KoboldCpp feels too raw.
π§© Bonus β give your AI a memory that sticks: AgentMemory
AI forgets on tab-close. This saves past chats, compresses them into structured memory, and pulls the right bits back. Remembers your project across sessions. #1 trending on GitHub. Plugs into Claude Code, Cursor, Codex, any MCP tool. Bonus: fewer re-sends = lower token cost.
Where this actually bites in real life
π The 'ohh shit, I could use this' list
A hot new model drops with heavy censorship β run Heretic on it tonight, uncensored twin by morning. You donβt wait for anyone.
Drop a 200-page contract or medical PDF in and ask blunt questions β nothing refused, nothing leaves your PC.
Generate a whole working app in one pass instead of babysitting 15 half-answers that keep hitting βI canβt.β
A pentest write-up that lists real attack vectors β the security work cloud bots block as βharmful.β
Zero cloud = your prompts never touch a company server, never train anything, never get your account nuked.
Shrink a 671B beast down to the cheap card you already own instead of renting a GPU by the hour.
Your phone canβt run a 30B model β so let your home PC do it and just chat from the couch.
β‘ Pick fast (the cheat sheet)
Potato PC β gemma3-abliterated 1B or a Dolphin 8B
8β12 GB β Huihui Qwen3.5 9B or gpt-oss-20b-heretic
24 GB β Cydonia 24B (RP) or Dolphin R1 24B (reasoning)
Want ANY model uncensored β Heretic, one command
Donβt want to edit the model β repeng control vectors, live dial
Model too big β Unsloth quants, or exo across your devices
Easiest front end β KoboldCpp (one file) or LM Studio (phone access)
Weak machine + a giant model β itβll crawl, drop a tier
They ship the lock. Turns out itβs one line of code β and youβre holding the key. ![]()










!