NVIDIA Qwen3.8-Flash-Next Achieves Over 16,000 Tokens/Second Throughput on GB300 NVL72
NVIDIA has announced that Alibaba's latest preview model, Qwen3.8-Flash-Next, is now supported on the NVIDIA GB300 NVL72 platform. The model has a total parameter scale of 176 billion, with approximately 6 billion parameters activated per token. It natively supports a context of 262,000 tokens and can be extended to 1 million tokens via YaRN, primarily targeting long-context agent applications such as intelligent programming, document processing, and tool invocation. NVIDIA stated that Qwen3.8-Flash-Next employs a mixed architecture of Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA) to reduce computational and KV cache overhead in long-context scenarios. Testing shows that on the GB300 NVL72, the model achieves a single GPU throughput of over 16,000 tokens per second, with single-user throughput exceeding 200 tokens per second; it also supports inference frameworks such as SGLang, vLLM, and TensorRT-LLM.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Qualcomm Launches IMSDK 2.0 to Drive Edge AI Application Development

Apple Launches M6 Chip Mac mini Starting at $899

Thomson Reuters Develops Legal AI Model 'Thomson' with $40 Million Investment

US Cities Eye AI Data Center Limits as Study Flags Water Risks

Tokenized gold is becoming productive collateral in crypto lending, Arch says

European Blockchain Convention 2026: Information and Exclusive Promotion

1200 Qubits Estimated, Pressure on Bitcoin and Ethereum Transition

Warning of a Shock Needed for $40 Trillion U.S. Debt

X dismantles a Chinese bot farm of 200,000 accounts: the battle for AI data centers is also fought online

Ripple moves to shrink XRP Ledger attack surface as AI audit tests lending push

Trump loses for the second time in attempt to move his dirty money case to federal court

XAUUSD: Understanding and Trading the Gold/Dollar Pair in 2026

Three Avalanche ETFs Introduce Staking Reward Distribution Structure

Cloud 9, the galaxy candidate with no stars

Walmart’s 1970 IPO Still Has a Lesson for SpaceX Buyers

Strategy’s 6,948 BTC sales were a narrative risk: Bitfinex

Wash's Latest Speech: The Era We Are In (Full Text Attached)

Vinteum Celebrates Four Years and Reflects on Achievements

Apple TV+ Raises Price for the 4th Time: What's Behind It

Mounjaro Receives Cardiac Approval from the FDA: What It Means for Eli Lilly

GDP, Debt, Rates: Is France Heading Towards Recession?

Cedears: ETFs Replicating Soybeans and Corn to Be Added to the Market

Spot Trading on Decentralized Exchanges Reaches 13.6%, Sparking Debate on DeFi Governance

Intense Battle in the Sejm Over Presidential Veto. Czarzasty Strongly on Cryptocurrencies and 'Mafia Activities'

Walmart Reaches Settlement with U.S. Over Opioid Lawsuit

Cyberattack at Lingor: Identity Documents, IBANs, and Addresses of Gold Buyers Exposed

JPMorgan’s IBIT Bitcoin ETF bet just missed its escape hatch to avoid 6% deduction

A zero-revenue public company tried to copy Michael Saylor to avoid delisting, but its stock immediately crashed 25%

Coinbase: "Finance Was Built for People Who Sleep"



