Install
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
openbmb/MiniCPM5-2B — 2B · DENSE · 128K ctx
1+ hour, 1+ min ago (167+ words) MiniCPM5-2B — dense 2B LLM with hybrid Think/No-Think reasoning, native 128K context, and native tool calling support, built on the standard Llama architecture 2B-class open-source model with strong reasoning and tool use MiniCPM5-2B is the 2B checkpoint in OpenBMB's MiniCPM5 series — a dense model built for…...
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp — 285B / 13B active · MOE · 1024K ctx
6+ day, 4+ hour ago (185+ words) DeepSeek's first experimental multimodal V4 model — the V4-Flash MoE backbone plus a 32-layer vision tower, 1M context, and a fused DSpark draft module. First multimodal V4 — 285B/13B MoE scoring 83.5% on OCRBench from a single GB200 NVL4 tray Vision support is not in a stable vLLM release…...
Qwen/Qwen3.8-Flash-Next — 176B / 6B active · MOE · 256K ctx
1+ week, 5+ day ago (331+ words) Qwen4 architecture preview with a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. Qwen4 architecture preview with 6B active parameters and efficient 262K context Qwen3.8-Flash-Next is a multimodal, ultra-sparse Mixture-of-Experts model. It has 125B parameters, including an additional 51B N-gram…...
Qwen/Qwen3.8-27B — 27B · DENSE · 256K ctx
3+ week, 2+ day ago (222+ words) 27B-parameter dense hybrid-attention model with linear attention on 48 of 64 layers, a vision tower, a built-in MTP draft head, 262K native context window and extensible to 1M context Fits one Blackwell GPU in every precision: NVFP4 in 24.6 GiB, 6.6M KV tokens at 1M context Qwen3.8-27B is the…...