heretic p-e-w

Fully automatic censorship removal for language models

Hauptsprache
Python
Stars
28,688
Forks
3,163

Tags

  • CLI-Tool
  • Deep Learning
  • GPU-Training
  • Forschung
  • Python

Zusammenfassung

Heretic entfernt Zensur (Safety-Alignment) aus transformer-basierten Sprachmodellen, ohne teures Post-Training. Es kombiniert eine fortschrittliche Implementierung von directional ablation ("Abliteration") mit einem TPE-basierten Parameter-Optimierer (Optuna) für vollständig automatische Optimierun…

README

<img width="128" align="right" alt="Logo" src="https://github.com/user-attachments/assets/df5f2840-2f92-4991-aa57-252747d7182e" /> # Heretic: Vollautomatische Entfernung von Zensur für Sprachmodelle<br><br>[![Discord](https://img.shields.io/discord/1447831134212984903?color=5865F2&label=discord&lab…

查看完整页面