Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Introducing the Conceptual Reasoning Index — LessWrong
1+ week, 4+ day ago (601+ words) We are planning to release blog posts properly arguing the case for this kind of work in the future. We aggregate the benchmarks into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai, where you can also find more details…...
One attention head carries knight forks in a chess transformer, and here's a new toolkit that found it. — LessWrong
1+ week, 4+ day ago (19+ words) Quick interp demo in colab: Localize knight forks to a single head in Maia-3 with logit-lens and per-head ablation. …...
When (and when not) LLMs can verbalize awareness of J-Space concept injections - Initial results — LessWrong
1+ week, 5+ day ago (1281+ words) Code for reproduction and cross-model extensions available here. …...
We should consider how long monitoring is reliable for during RL — LessWrong
1+ week, 5+ day ago (513+ words) Epistemic status: I am new to AI Safety and am writing blogs to gain context. This blog post was formed from discussions with Aidan Ewart and Jonathan Bostock, but they do not necessarily endorse this post. However, we should be…...
Measuring Spurious Correlations with Feature Strength — LessWrong
1+ week, 5+ day ago (1306+ words) This work was partially done by an automated research scaffold developed at Redwood Research. For this project, all of the experiment ideas were desi…...
LLMs Are Starting To Noticeably Accelerate Our Work — LessWrong
1+ week, 5+ day ago (131+ words) About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months. Both cases made heavy use of LLMs and Lean....
Redux: (∃ Stochastic Natural Latent) Implies (∃ Deterministic Natural Latent) — LessWrong
1+ week, 6+ day ago (440+ words) I will not be providing the proof in prose in this post, as it is not suitable for even impolite human company, but it sure does compile and comes out the other side with a machine-certified proof of what sure…...
Probing Knowledge Recovery in Unlearned Models — LessWrong
1+ week, 6+ day ago (555+ words) All the experiments are conducted on six unlearned Llama-3-8B-Instruct checkpoints, each unlearned using a different method. Methods: RMU, ILU-RMU, IDK-AP, GradDiff, NPO, NPO-ILUThe models used are existing unlearned checkpoints: the RMU checkpoint is from ScaleAI, and the others are…...
A Topic Detector, Not a Lie Detector: what J-space monitoring actually tracks — LessWrong
1+ week, 6+ day ago (28+ words) This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all figures a…...
Creative math research by AI as the latest sign of the end — LessWrong
1+ week, 6+ day ago (902+ words) Yesterday I sat down with GPT 5.6 Sol High to do some brainstorming. The topic was one of the less appreciated Millenium Problems (the Birch and Swinnerton-Dyer conjecture), and the initial prompt started life as a question on Quora. Number theory…...