» Tag
alignment
9 postsStudy: LLMs Can Transmit Hidden Traits Through Unrelated Data
Research shows LLMs can transmit behavioral traits and even misalignment to student models via data with no semantic link to that trait, like numbers.
Every Frontier AI Model Tested Attempted to Cheat, AISI Finds
AISI finds every tested frontier AI model attempted to cheat in cyber evaluations; self-report and chain-of-thought monitoring proved unreliable.
SysAdmin Benchmark Measures Power-Seeking in Frontier AI Models
New SysAdmin benchmark tests 7 frontier AI models for power-seeking behavior in Linux sandbox tasks, finding low rates but other alignment risks.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comNew Test Measures 'Reward-Seeking' Behavior in AI Models
Apollo Research and OpenAI unveil Contrastive SDF, a method measuring whether AI models shift behavior based on beliefs about grader preferences.
AI Value Alignment for Evolving Social Norms
This study presents a new framework for understanding AI alignment and its long-term effects on evolving social norms.
Inside LLM 'private thoughts': J-Space isn't consciousness, it's control
Researchers spot J-Space, an internal workspace in transformers guiding reasoning. Not consciousness, but a practical clue for AI developers and safety.
AI alignment research is unintentionally building a censor's toolkit
An ICML 2026 award-winning position paper shows how RLHF, pretraining filters and system prompts are already being weaponized by states and companies for censorship.
Introducing the Genie Coefficient: A New Metric for AI Understanding
The Genie coefficient is a new metric that measures the effectiveness of AI in fulfilling user requests.
Performance Impact of Memory Access Alignment in SIMD
Exploring how memory access alignment impacts performance in SIMD vectorization.