Install
The New Stack is a media platform for the people who build and manage software the world relies on. We provide context and explanation of at-scale technologies to advance knowledge and create conversations through our coverage of modern architectures, components of the software development life cycle, and operations to
- 152articles · 30d
- 8+ hour agolatest article
- Aug 14, 2026earliest in window
- 96%with images
- 88avg words
- Science & Technology 140
- Software Dev. 101
- Computers & Electronics 95
- News 38
- Software 21
- Science & Nature 13
- Economy, Business & Finance 9
- Finance & Business 9
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.
4+ day, 3+ hour ago (492+ words) Sierra has open-sourced Hyper-𝜏-bench, a follow-up to its 2024 τ-bench that tests how well AI agents can build other agents....
K2 Horizon just shipped as six new fully open models — developers aren't fully convinced
4+ day, 11+ hour ago (632+ words) Based in the Emirati capital, Abu Dhabi, the Institute of Foundation Models (IFM) introduced K2 Horizon last week. This group of six AI foundation models, ranging from 0.9 billion to 375 billion parameters, is claimed to be the “largest fully open-source fleet of…...
GLM-5.3-Flash vs. GLM-5.3: Time and money, not the spec sheet
1+ week, 5+ day ago (801+ words) I always wonder about the end goal when companies launch products so close together and undercut each other by claiming the new one is “so much better.” GLM-5.3 and GLM-5.3-Flash are a great example of this. Z.AI launched…...
DeepSeek's first vision model vs. Gemini 3.7 Flash: It comes down to spend vs. speed
1+ week, 6+ day ago (334+ words) DeepSeek matched Gemini on nine vision questions for a third of the price. The tradeoff: slower, less consistent responses that matter at scale....
IBM's new Granite 4.2 models add reasoning and stay dense
2+ week, 5+ day ago (50+ words) IBM’s Granite 4.2 sticks with decoder-only models, adds a 512,000-token context window, and trains its larger versions for agentic work....
An industrial-scale distillation of models, or subtle benchmaxxing: What developers really think of GLM-5.3
3+ week, 4+ day ago (536+ words) Z.ai's GLM-5.3 claims big coding gains, but AI professionals say the real story may be distillation from Anthropic's Claude models, or a case of subtle bencmaxxing, not novel training....