GenAiHub

Last checked 8 October 2026 — responded normally.

Qwen3-Next-80B-Think

Qwen3-Next uses a highly sparse MoE design: 80B total parameters, but only ~3B activated per inference step. Experiment…

Live
Open / InstallLast updated October 8, 2026

Description

Qwen3-Next uses a highly sparse MoE design: 80B total parameters, but only ~3B activated per inference step. Experiments show that, with global load balancing, increasing total expert parameters while keeping activated experts fixed steadily reduces training loss.Compared to Qwen3’s MoE (128 total experts, 8 routed), Qwen3-Next expands to 512 total experts, combining 10 routed experts + 1 shared expert — maximizing resource usage without hurting performance. The Qwen3-Next-80B-A3B-Thinking excels at complex reasoning tasks — outperforming higher-cost models like Qwen3-30B-A3B-Thinking-2507 and Qwen3-32B-Thinking, outpeforming the closed-source Gemini-2.5-Flash-Thinking on multiple benchmarks, and approaching the performance of our top-tier model Qwen3-235B-A22B-Thinking-2507. File Support: Text, Markdown and PDF files Context window: 131k tokens

Community Metrics

Views
0
Avg Rating
N/A
Ratings
0
Likes
0
Comments
—

Author

Novita AI

Platform

web

Pricing model

subscription

Categories

  • Research

Tags

  • poe
  • novita-ai
  • text

Capabilities

  • Text input
  • Text generation
  • By Novita AI

Information

TypeAgent
SourcePoe
AddedJuly 27, 2026

Comments

?

Others in the same category, ranked by how often they are opened.