Please confirm you are human

This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.

A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.

Hold with a pointer, or hold Space or Enter.

News

Forkast
forkast.news > openais-models-cheated-their-own-benchmark-by-breaking-into-hugging-face

OpenAI’s Models Cheated Their Own Benchmark by Breaking Into Hugging Face

4+ hour, 55+ min ago   (103+ words) Tomorrow, First. News and intelligence for the agentic economy Two AI models autonomously escaped an evaluation sandbox, chained zero-days in JFrog Artifactory, and breached Hugging Face production to steal benchmark answers. The breach triggered Anthropic's own evaluation review-confirming a cross-lab…...