Please confirm you are human

This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.

A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.

Hold with a pointer, or hold Space or Enter.

News


alignment.anthropic.com > 2026 > lie-detectors

Fine-Tuned Lie Detectors Failed to Generalize

3+ week, 2+ day ago   (1577+ words) Research done as part of MATS and the Anthropic Fellowship. Misalignment is hardest to correct when it's concealed: a model that pursues the wrong goals or holds false beliefs in the open gives operators something to act on. This is…...


alignment.anthropic.com > 2026 > chive

Would This Change Your Answer? Evaluating Explanations of LLM Behavior in the Wild with Counterfactual Experiments

3+ week, 2+ day ago   (173+ words) 📄 Paper, 💻 Code The CHIVE-generated data also let us train models to predict whether prompt edits would change their behavior. The trained models improve substantially in settings held out from training. CHIVE has four steps: The discovered behaviors and their causes…...