Ant Group Launches AI Model with 124 Billion Parameters Surpassing Its One Trillion Giant
Ant Group surprises the AI world with Ling-3.0-Flash, an open-weight model with 124 billion parameters that activates only 5.1 billion per token and surpasses benchmarks of its one trillion predecessor. A feat that redefines computational efficiency and the cost of implementing AI agents.
- Ling-3.0-Flash uses a MoE architecture with a total of 124B parameters and only 5.1B active per token, achieving parity or surpassing Ling-2.6-1T on key tasks.
- The model is available under MIT license and features a native context window of 262,000 tokens, aiming for one million, designed for high-frequency production agents.
Ant Group, the fintech giant behind Alipay, launched Ling-3.0-Flash on July 23, 2026. The model was developed by its inclusionAI lab and challenges the notion that in AI, bigger is always better.
Ling-3.0-Flash has a total of 124 billion parameters but activates only around 5.1 billion per token during inference. This Mixture of Experts (MoE) architecture allows only a subset of the neural network to engage for each input, details Cryptobriefing.
Performance That Challenges Scale
In the AI Intelligence Analysis Index, the model scored 38, matching or surpassing its predecessor Ling-2.6-1T, which had a trillion complete parameters. This makes it approximately eight times smaller in total parameter count but competitive in the metrics that matter for production deployments.
The main achievement lies in core reasoning and instruction-following tasks. The architecture combines hybrid reasoning with MoE, achieving efficiency without sacrificing quality.
Context length also advances: the native window is 262,000 tokens, with plans to expand to one million. The attention mechanism combines Kimi’s Delta Attention layers and Multi-Head Latent Attention, known as a hybrid-linear approach.
This design keeps memory and computation requirements manageable as context grows. Ant Group describes this technique as fundamental to supporting complex agent tasks.
Designed for Agents, Not Just for Chat
Ling-3.0-Flash was specifically designed for production-grade AI agents running at high frequency. These workflows demand rapid token generation, reliable adherence to instructions, and handling of long context without degradation.
A massive library where the librarian only needs to pull five books at a time, regardless of how complex the question is. That’s the metaphor developers use to explain the approach.
Inference efficiency is key to reducing costs. By activating only 5.1B parameters, the computation required per query is directly reduced, making operation cheaper at scale.
Ant Group has already tested the model in its own Alipay services. Although specific figures were not disclosed, the company emphasizes that the model is optimized to handle complex production agent tasks.
Open and Commercial Availability
The model was published on Hugging Face under the MIT license, the most permissive within open source. This allows for use, modification, and commercial redistribution without royalties.
Free access to the API was also offered through OpenRouter and Kilo until August 3, 2026. Additionally, it is accessible through Ant Group’s channels and Vercel’s AI gateway.
The MIT license removes barriers for developers and companies wanting to integrate it into their products. This contrasts with other open-source models that impose restrictions on commercial use.
Ant Group’s decision to open the model aims to position the company as a leader in efficiency. According to the original report, this could accelerate AI adoption in sectors where inference costs are prohibitive.
Implications for the Future of AI
If a 124B MoE model with 5.1B active can match or surpass the baseline of a trillion parameters, the cost structure for implementing capable AI decreases significantly. Inference is where AI companies spend money after training.
This advancement could pressure other labs to rethink their scaling strategies. Efficiency becomes a key competitive factor, not just the number of parameters.
The AI agent market directly benefits, as response times are critical. Models like Ling-3.0-Flash allow for faster interactions and predictable costs.
CryptoBriefing highlights that this is a notable step in the AI efficiency race, with implications for the entire industry. The combination of speed, low cost, and open license could change the landscape of artificial intelligence.
Ant Group demonstrates that innovation does not always mean building the largest model, but the smartest for real-world use. The development community can already experiment with this new generation of efficient AI.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Who Oversees When AI Agents Hire Each Other?

Ethereum Launches Testnet Platåberget for Glamsterdam

Curve founder says FATF pressure could make DeFi safer and more decentralized

Turkish Investors to Start Graphite Mining in Ukraine's Khmelnytskyi Region

Libya Seeks $40 Billion Investment in Oil and Gas Sector

Iranian Speaker Says Strait of Hormuz Will Not Open Until Conditions Are Met

Unitree Crypto Futures Valued at $40 Billion, IPO at $9 Billion

BitMart Exchange Closure: Will Sheldon Lee Respond to the Ultimatum by August 19?

Asseto Launches PGNGI+ Introducing Equity from United Group's Private Infrastructure Fund

Wyoming Blockchain Summit Opens: The Focus of the Crypto Industry Shifts from Product Licensing to Market Structure

Bitcoin Core v32 Enters Feature Freeze on August 20

OpenAI President Warns of AI Security Risks and Outlines 10 Countermeasures

YGG 3.0 Officially Launched, Transitioning to Infrastructure Layer

Bitcoin Supply on Exchanges Rose to 1.332 Million BTC in August

Cyberattack in Taiwan Uses AI Against Government Agencies
HSBC: Fed's Reduction of Holdings Will Make US Treasuries More Dependent on Price-Sensitive Buyers

Ethereum's next major upgrade just slipped to late 2026, forcing a two-week scramble to save its 2027 roadmap

Ledger Challenges the Crypto Sector: Security in Bitcoin Design is Non-Negotiable

G10 Forex Low Volatility Expected to Persist Until Early September

What is the relationship between BONK and Solana? Understanding BONK's position in the Solana ecosystem

Fed's $17 Billion Purchase Not New Quantitative Easing

30% Chance of Rate Hike in September as Emerging Market Currencies Rebound

Anchorpoint Announces Parallel Operation of On-chain and Off-chain Control Measures for HKDAP

People's Bank of China Adds 8 New Digital Renminbi Operating Institutions

web3: Foreign Media: High Bitcoin Futures Positions May Amplify Volatility

Crypto perpetuals surge from $15 billion to $250 billion in three months, according to CryptoQuant

BitMart Employees Speak Out: Demand Disclosure of Platform Assets and Repayment Plan by August 19

ZODL Completes Migration of Zodl Wallet from Orchard to Ironwood, Fixes Zebra Backend Issues






