Add Weibo (Sina Weibo) as a lab? — VibeThinker reasoning models

#43 · open · 0 comments

View on GitHub ↗

hammer

## Proposal: add **Weibo (Sina Weibo)** as a lab, with VibeThinker as its output Research + recommendation per the file-an-issue-before-big-changes convention. Neither Weibo nor VibeThinker is currently tracked. ### The artifact — VibeThinker ([arXiv 2606.16140](https://arxiv.org/abs/2606.16140), 2026-06-15) - **VibeThinker-3B** + **VibeThinker-1.5B** — compact dense reasoning models from **Sina Weibo Inc.** (corresponding authors `@staff.weibo.com`; HF [`WeiboAI`](https://huggingface.co/WeiboAI), GitHub [`WeiboAI/VibeThinker`](https://github.com/WeiboAI/VibeThinker)). **MIT** license. - **Post-train of `Qwen2.5-Coder-3B`** (not a from-scratch foundation model). Novel recipe: **"Spectrum-to-Signal" / diversity-driven optimization** — curriculum SFT → multi-domain RL → offline self-distillation. - **Self-reported** benchmarks: AIME26 **94.3** (97.1 w/ test-time scaling), LiveCodeBench v6 **80.2** pass@1, 96.1% acceptance on unseen LeetCode, IFEval 93.4 — claimed to match/exceed DeepSeek-V3.2, GLM-5, Gemini 3 Pro on verifiable reasoning *at orders of magnitude fewer params*. - **Traction:** GitHub ★859, HF ♥532 (1.5B) / ♥280 (3B) within days; [VentureBeat coverage](https://venturebeat.com/technology/why-weibos-tiny-vibethinker-3b-has-the-ai-world-arguing-over-benchmarks-again). ### On-thesis assessment — **yes** Squarely in scope on two axes: (1) frontier **reasoning** (small-model regime) and (2) a **foundational post-training technique** (diversity-driven optimization eliciting large-model reasoning in a tiny model). Open weights, real attention, from a major company. ### Adversarial review (the case against / caveats) 1. **Single release, Qwen fine-tune.** `WeiboAI` is a brand-new HF org with exactly one release (3B + 1.5B), one GitHub repo, one paper. It's a post-train of Qwen2.5-Coder, not a pretrained FM. Adding a *lab* on one output risks a one-off. 2. **The headline claim is self-reported and domain-narrow.** **Not scored on Artificial Analysis** (no independent composite). The "matches Gemini 3 Pro" framing holds only for *verifiable* reasoning — on **GPQA-Diamond it scores 70.2 vs Gemini 3 Pro's 91.9 / Opus 4.5's 87.0**. The release [triggered a public benchmark dispute](https://venturebeat.com/technology/why-weibos-tiny-vibethinker-3b-has-the-ai-world-arguing-over-benchmarks-again); community verdict: "legitimate but domain-specific." 3. **No prior Weibo AI-research track record** in our data — Weibo is a social-media platform, not historically a model lab. ### Recommendation — **ADD, lightweight, with accurate/caveated recording** The technique is genuinely novel, the weights are MIT-open, and the result is notable enough to be debated industry-wide — that clears our inclusion bar, and we already track comparable single/few-output entrants (Liquid AI, Reka, Inception Labs, Sarvam). Concretely: - New lab `weibo` (name "Weibo" / Sina Weibo), region china, type public, framed as an **emerging** AI entrant. - One output `vibethinker` (type: model), base_model Qwen2.5-Coder-3B, MIT, with **self-reported** `benchmark_scores`, the GPQA-Diamond caveat in the description, and **`intelligence_index` left unset** (AA hasn't scored it — avoids overstating the "frontier" claim). - Tags: reasoning, post-training, open-weight, small-model. - Revisit lab depth if Weibo ships a second substantive release. **Alternative considered:** wait for a second release before adding a lab. Reasonable, but the technique + open weights + attention make it worth tracking now, and outputs can't currently be lab-less (see #11), so tracking VibeThinker requires the lab. Holding implementation pending a 👍 on the recommendation.

Comments