heretic p-e-w
Fully automatic censorship removal for language models
- Hauptsprache
- Python
- Stars
- 28,688
- Forks
- 3,163
Tags
- CLI-Tool
- Deep Learning
- GPU-Training
- Forschung
- Python
Zusammenfassung
Heretic entfernt Zensur (Safety-Alignment) aus transformer-basierten Sprachmodellen, ohne teures Post-Training. Es kombiniert eine fortschrittliche Implementierung von directional ablation ("Abliteration") mit einem TPE-basierten Parameter-Optimierer (Optuna) für vollständig automatische Optimierun…
README
<img width="128" align="right" alt="Logo" src="https://github.com/user-attachments/assets/df5f2840-2f92-4991-aa57-252747d7182e" /> # Heretic: Vollautomatische Entfernung von Zensur für Sprachmodelle<br><br>[![Discord](https://img.shields.io/discord/1447831134212984903?color=5865F2&label=discord&lab…