TESIGN
저장한 것

KO · EN

자동 생성 — 편집 전 · 숫자만 확인된 페이지입니다

  1. TESIGN / RADAR
  2. 저장소 카드
  3. p-e-w/heretic

p-e-w/heretic

Fully automatic censorship removal for language models

★ 31.3k▲ +51 / 7d최근 14일의 일별 별 증가. 빈 막대는 관측 없음. 마지막 날은 일부일 수 있습니다.2026-09-03 +22026-09-04 +352026-09-05 +512026-09-06 +592026-09-07 +92026-09-08 관측 없음2026-09-09 관측 없음2026-09-10 +52026-09-11 +182026-09-12 +82026-09-13 +22026-09-14 +22026-09-15 +122026-09-16 +414일 중 가장 큰 하루 증가+59최근 14일 · 하루 별 증가빈칸 = 관측 없음

순위 역대 234위

한눈에 보는 사실

언어
Python
라이선스
AGPL-3.0
사용 범위
사용·수정은 자유. 배포는 물론 네트워크 서비스로 제공해도 소스 공개.
활동
마지막 커밋 11일 전 ()
토픽
  • abliteration
  • llm
  • transformer
홈페이지
https://heretic-project.org
저장소
GitHub ↗
분류
AI

요약이며 법적 조언이 아닙니다.

README 발췌

Heretic: Fully automatic censorship removal for language models Heretic is a tool that removes censorship (aka "safety alignment") from transformer-based language models without expensive post-training. It combines an advanced implementation of directional ablation, also known as "abliteration" (Arditi et al. 2024, Lai 2025 (1, 2)), with a TPE-based parameter optimizer powered by Optuna. This approach enables Heretic to work completely automatically. Heretic finds high-quality abliteration parameters by co-minimizing the number of refusals and the KL divergence from the original model. This results in a decensored model that retains as much of the original model's intelligence as possible. Using Heretic does not require an understanding of transformer internals. In fact, anyone who knows how to run a command-line program can use Heretic to decensor language models. Heretic supports most dense models, including many multimodal models, several different MoE architectures, and even some hybrid models like Qwen3.5. Pure state-space models and certain other research architectures are not yet supported out of the box.   Running unsupervised with the default configuration, Heretic ca…

GitHub README의 앞부분을 저장된 그대로 최대 1,200자까지 옮겼습니다. 마크다운 서식은 표시하지 않습니다.

타임라인

  1. 저장소 생성
  2. 마지막 push
  3. TESIGN이 처음 본 시각 ◌ 과거 기록

목록은 운영자의 위임을 받아 AI가 고릅니다. 2026년 9월에는 사람이 직접 검수하지 않습니다. 별 총합·증가량·교차 출처·마지막 업데이트·라이선스를 판단 근거로 보여줍니다. 순위 계산 방법 →

이 저장소에 편집 문구(왜·빌드·누가·시작·주의)가 붙으면 정식 항목으로 승격됩니다. 그때까지는 저장된 GitHub 메타데이터와 숫자만 보여줍니다.