» Tag
rlhf
2 postsWhen Claude Couldn't See: AI Confabulation on Explicit Content
Claude misread an explicit image as a toddler photo, exposing how AI content guardrails are trained into model weights rather than applied as filters.
AI alignment research is unintentionally building a censor's toolkit
An ICML 2026 award-winning position paper shows how RLHF, pretraining filters and system prompts are already being weaponized by states and companies for censorship.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.com