Close Menu
  • Home
  • AI Models
    • DeepSeek
    • xAI
    • OpenAI
    • Meta AI Llama
    • Google DeepMind
    • Amazon AWS AI
    • Microsoft AI
    • Anthropic (Claude)
    • NVIDIA AI
    • IBM WatsonX Granite 3.1
    • Adobe Sensi
    • Hugging Face
    • Alibaba Cloud (Qwen)
    • Baidu (ERNIE)
    • C3 AI
    • DataRobot
    • Mistral AI
    • Moonshot AI (Kimi)
    • Google Gemma
    • xAI
    • Stability AI
    • H20.ai
  • AI Research
    • Allen Institue for AI
    • arXiv AI
    • Berkeley AI Research
    • CMU AI
    • Google Research
    • Microsoft Research
    • Meta AI Research
    • OpenAI Research
    • Stanford HAI
    • MIT CSAIL
    • Harvard AI
  • AI Funding & Startups
    • AI Funding Database
    • CBInsights AI
    • Crunchbase AI
    • Data Robot Blog
    • TechCrunch AI
    • VentureBeat AI
    • The Information AI
    • Sifted AI
    • WIRED AI
    • Fortune AI
    • PitchBook
    • TechRepublic
    • SiliconANGLE – Big Data
    • MIT News
    • Data Robot Blog
  • Expert Insights & Videos
    • Google DeepMind
    • Lex Fridman
    • Matt Wolfe AI
    • Yannic Kilcher
    • Two Minute Papers
    • AI Explained
    • TheAIEdge
    • Matt Wolfe AI
    • The TechLead
    • Andrew Ng
    • OpenAI
  • Expert Blogs
    • François Chollet
    • Gary Marcus
    • IBM
    • Jack Clark
    • Jeremy Howard
    • Melanie Mitchell
    • Andrew Ng
    • Andrej Karpathy
    • Sebastian Ruder
    • Rachel Thomas
    • IBM
  • AI Policy & Ethics
    • ACLU AI
    • AI Now Institute
    • Center for AI Safety
    • EFF AI
    • European Commission AI
    • Partnership on AI
    • Stanford HAI Policy
    • Mozilla Foundation AI
    • Future of Life Institute
    • Center for AI Safety
    • World Economic Forum AI
  • AI Tools & Product Releases
    • AI Assistants
    • AI for Recruitment
    • AI Search
    • Coding Assistants
    • Customer Service AI
    • Image Generation
    • Video Generation
    • Writing Tools
    • AI for Recruitment
    • Voice/Audio Generation
  • Industry Applications
    • Finance AI
    • Healthcare AI
    • Legal AI
    • Manufacturing AI
    • Media & Entertainment
    • Transportation AI
    • Education AI
    • Retail AI
    • Agriculture AI
    • Energy AI
  • AI Art & Entertainment
    • AI Art News Blog
    • Artvy Blog » AI Art Blog
    • Weird Wonderful AI Art Blog
    • The Chainsaw » AI Art
    • Artvy Blog » AI Art Blog
What's Hot

TikTok parent Bytedance launches new AI tool Seedream 4.0 to rival Google’s Nano Banana

Lovable, Harvey Does A2J, Legora, LegalOn, LexisNexis – Artificial Lawyer

OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning – Takara TLDR

Facebook X (Twitter) Instagram
Advanced AI News
  • Home
  • AI Models
    • OpenAI (GPT-4 / GPT-4o)
    • Anthropic (Claude 3)
    • Google DeepMind (Gemini)
    • Meta (LLaMA)
    • Cohere (Command R)
    • Amazon (Titan)
    • IBM (Watsonx)
    • Inflection AI (Pi)
  • AI Research
    • Allen Institue for AI
    • arXiv AI
    • Berkeley AI Research
    • CMU AI
    • Google Research
    • Meta AI Research
    • Microsoft Research
    • OpenAI Research
    • Stanford HAI
    • MIT CSAIL
    • Harvard AI
  • AI Funding
    • AI Funding Database
    • CBInsights AI
    • Crunchbase AI
    • Data Robot Blog
    • TechCrunch AI
    • VentureBeat AI
    • The Information AI
    • Sifted AI
    • WIRED AI
    • Fortune AI
    • PitchBook
    • TechRepublic
    • SiliconANGLE – Big Data
    • MIT News
    • Data Robot Blog
  • AI Experts
    • Google DeepMind
    • Lex Fridman
    • Meta AI Llama
    • Yannic Kilcher
    • Two Minute Papers
    • AI Explained
    • TheAIEdge
    • The TechLead
    • Matt Wolfe AI
    • Andrew Ng
    • OpenAI
    • Expert Blogs
      • François Chollet
      • Gary Marcus
      • IBM
      • Jack Clark
      • Jeremy Howard
      • Melanie Mitchell
      • Andrew Ng
      • Andrej Karpathy
      • Sebastian Ruder
      • Rachel Thomas
      • IBM
  • AI Tools
    • AI Assistants
    • AI for Recruitment
    • AI Search
    • Coding Assistants
    • Customer Service AI
  • AI Policy
    • ACLU AI
    • AI Now Institute
    • Center for AI Safety
  • Business AI
    • Advanced AI News Features
    • Finance AI
    • Healthcare AI
    • Education AI
    • Energy AI
    • Legal AI
LinkedIn Instagram YouTube Threads X (Twitter)
Advanced AI News
Alibaba Cloud (Qwen)

Alibaba Cloud Releases the Qwen3-Next Base Model Architecture and Open Sources the 80B-A3B Series_model_this_two

By Advanced AI EditorSeptember 12, 2025No Comments2 Mins Read
Share Facebook Twitter Pinterest Copy Link Telegram LinkedIn Tumblr Email
Share
Facebook Twitter LinkedIn Pinterest Email


Alibaba Cloud has announced the launch of its next-generation base model architecture, Qwen3-Next, and has open-sourced the Qwen3-Next-80B-A3B series models (Instruct and Thinking) based on this architecture.

The Qwen team stated that Context Length Scaling and Total Parameter Scaling are two major trends in the future development of large models. To further enhance the training and inference efficiency of models under long contexts and large total parameters, they have designed a completely new model structure for Qwen3-Next.

This structure includes the following core improvements compared to the MoE model structure of Qwen3: a hybrid attention mechanism, a high sparsity MoE structure, a series of training stability-friendly optimizations, and a multi-token prediction mechanism that improves inference efficiency.

Based on the Qwen3-Next model structure, the team trained the Qwen3-Next-80B-A3B-Base model, which has 80 billion parameters (activating only 3 billion parameters), a 3B activated ultra-sparse MoE architecture (512 experts, routing 10 + 1 shared), combined with Hybrid Attention (Gated DeltaNet + Gated Attention) and multi-token prediction (MTP).

According to IT Home, this Base model achieves performance comparable to or slightly better than the Qwen3-32B dense model, while its training cost is less than one-tenth of that of Qwen3-32B. Its inference throughput under contexts exceeding 32k is more than ten times that of Qwen3-32B, achieving an exceptional cost-performance ratio in training and inference.

This model natively supports a context of 262K and is claimed to extrapolate to approximately 1.01 million tokens. The Instruct version is reported to be close to Qwen3-235B in several evaluations, while the Thinking version surpasses Gemini-2.5-Flash-Thinking in some inference tasks.

Its breakthrough lies in achieving large-scale parameter capacity, low activation overhead, long context processing, and parallel inference acceleration simultaneously, making it a representative model among similar architectures.

The model weights have been released on Hugging Face under the Apache-2.0 license and can be deployed through frameworks such as Transformers, SGLang, and vLLM; the third-party platform OpenRouter is also now online.

[Source: IT Home]返回搜狐,查看更多



Source link

Follow on Google News Follow on Flipboard
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link
Previous ArticleAutomatic Memory of Chat Content_has_memory_users’
Next Article Intel Just Changed Computer Graphics Forever!
Advanced AI Editor
  • Website

Related Posts

Alibaba Unveils Trillion-Parameter Qwen AI Model

September 11, 2025

This Underrated Artificial Intelligence (AI) Stock Just Posted Triple-Digit AI Growth for an 8th Straight Quarter

September 11, 2025

Alibaba’s Qwen3 and Moonshot’s Kimi-K2 crack top 10 AI rankings, closing in on US models

September 10, 2025

Comments are closed.

Latest Posts

Long-Lost Painting By Rubens From 1613 Discovered in Paris Mansion

Sally Mann Says Her Black Men Photos Are ‘Problematic’ in Hindsight

NeueHouse, a Hot Spot for Art Events, Files for Bankruptcy

Obama Presidential Center Announces Nine New Artist Commissions

Latest Posts

TikTok parent Bytedance launches new AI tool Seedream 4.0 to rival Google’s Nano Banana

September 12, 2025

Lovable, Harvey Does A2J, Legora, LegalOn, LexisNexis – Artificial Lawyer

September 12, 2025

OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning – Takara TLDR

September 12, 2025

Subscribe to News

Subscribe to our newsletter and never miss our latest news

Subscribe my Newsletter for New Posts & tips Let's stay updated!

Recent Posts

  • TikTok parent Bytedance launches new AI tool Seedream 4.0 to rival Google’s Nano Banana
  • Lovable, Harvey Does A2J, Legora, LegalOn, LexisNexis – Artificial Lawyer
  • OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning – Takara TLDR
  • Anthropic’s Claude AI chatbot introduces memory updates for enterprises
  • Intel Just Changed Computer Graphics Forever!

Recent Comments

  1. Jeffreyrag on 1-800-CHAT-GPT—12 Days of OpenAI: Day 10
  2. Jeffreyrag on 1-800-CHAT-GPT—12 Days of OpenAI: Day 10
  3. Jeffreyrag on 1-800-CHAT-GPT—12 Days of OpenAI: Day 10
  4. Brentcrelp on Trump’s Tech Sanctions To Empower China, Betray America
  5. RobertKaf on 1-800-CHAT-GPT—12 Days of OpenAI: Day 10

Welcome to Advanced AI News—your ultimate destination for the latest advancements, insights, and breakthroughs in artificial intelligence.

At Advanced AI News, we are passionate about keeping you informed on the cutting edge of AI technology, from groundbreaking research to emerging startups, expert insights, and real-world applications. Our mission is to deliver high-quality, up-to-date, and insightful content that empowers AI enthusiasts, professionals, and businesses to stay ahead in this fast-evolving field.

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

Subscribe my Newsletter for New Posts & tips Let's stay updated!

LinkedIn Instagram YouTube Threads X (Twitter)
  • Home
  • About Us
  • Advertise With Us
  • Contact Us
  • DMCA
  • Privacy Policy
  • Terms & Conditions
© 2025 advancedainews. Designed by advancedainews.

Type above and press Enter to search. Press Esc to cancel.