자동 생성 — 편집 전 · 숫자만 확인된 페이지입니다
- TESIGN / RADAR
- 저장소 카드
- p-e-w/heretic
p-e-w/heretic
Fully automatic censorship removal for language models
순위 역대 234위
한눈에 보는 사실
- 언어
- Python
- 라이선스
- AGPL-3.0
- 사용 범위
- 사용·수정은 자유. 배포는 물론 네트워크 서비스로 제공해도 소스 공개.
- 활동
- 마지막 커밋 11일 전 ()
- 토픽
- abliteration
- llm
- transformer
- 홈페이지
- https://heretic-project.org
- 저장소
- GitHub ↗
- 분류
- AI
요약이며 법적 조언이 아닙니다.
README 발췌
Heretic: Fully automatic censorship removal for language models Heretic is a tool that removes censorship (aka "safety alignment") from transformer-based language models without expensive post-training. It combines an advanced implementation of directional ablation, also known as "abliteration" (Arditi et al. 2024, Lai 2025 (1, 2)), with a TPE-based parameter optimizer powered by Optuna. This approach enables Heretic to work completely automatically. Heretic finds high-quality abliteration parameters by co-minimizing the number of refusals and the KL divergence from the original model. This results in a decensored model that retains as much of the original model's intelligence as possible. Using Heretic does not require an understanding of transformer internals. In fact, anyone who knows how to run a command-line program can use Heretic to decensor language models. Heretic supports most dense models, including many multimodal models, several different MoE architectures, and even some hybrid models like Qwen3.5. Pure state-space models and certain other research architectures are not yet supported out of the box. Running unsupervised with the default configuration, Heretic ca…
GitHub README의 앞부분을 저장된 그대로 최대 1,200자까지 옮겼습니다. 마크다운 서식은 표시하지 않습니다.
타임라인
- 저장소 생성
- 마지막 push
- TESIGN이 처음 본 시각 ◌ 과거 기록
목록은 운영자의 위임을 받아 AI가 고릅니다. 2026년 9월에는 사람이 직접 검수하지 않습니다. 별 총합·증가량·교차 출처·마지막 업데이트·라이선스를 판단 근거로 보여줍니다. 순위 계산 방법 →
이 저장소에 편집 문구(왜·빌드·누가·시작·주의)가 붙으면 정식 항목으로 승격됩니다. 그때까지는 저장된 GitHub 메타데이터와 숫자만 보여줍니다.