Ant Bailing Open-Sources 7.9B-Parameter Hybrid Reasoning MoE Model Ling-3.0-tiny
Ant Group's Bailing AI unit has open-sourced Ling-3.0-tiny, a lightweight hybrid reasoning MoE model with 7.9 billion total parameters and 1.3 billion activated during inference. Available in BF16, FP8 and INT4 versions, it scores 25 on the Artificial Analysis Intelligence Index, one point below Gemma-4-26B-A4B. The model pairs KDA and MLA attention in a 3:1 ratio with 128 routing experts, supports native hybrid reasoning, and achieves 100-105 tokens per second on DGX Spark at FP8 precision.
Ant Bailing has open-sourced Ling-3.0-tiny, a lightweight hybrid reasoning mixture-of-experts (MoE) model. The model has 7.9 billion total parameters with 1.3 billion activated during inference, and is designed primarily for efficient local deployment. It is offered in BF16, FP8 and INT4 versions, and has been verified on DGX Spark, MacBook and Mac mini devices. On the Artificial Analysis Intelligence Index, Ling-3.0-tiny scored 25, one point below Gemma-4-26B-A4B and above gpt-oss-120B (high), Qwen3.5-9B, Gemma-4-12B and Gemma-4-E4B.
Ling-3.0-tiny uses Moonshot AI's self-developed KDA (Kimi Incremental Attention) and DeepSeek's MLA (Multi-head Latent Attention) in a 3:1 alternating stack, paired with a sparse MoE feed-forward network comprising 128 routing experts. Each token activates only 8 routing experts and 1 shared expert, achieving a balance between long-context modeling capability, parameter efficiency and computational cost. The model supports native hybrid reasoning: the enable_thinking parameter allows thinking mode to be flexibly toggled on or off within a single request. It is suited to general agent tasks, code generation, mathematical and scientific reasoning, and instruction-following scenarios.
On deployment, at FP8 precision Ling-3.0-tiny delivers inference throughput of 100-105 tokens per second on DGX Spark and 86-90 tokens per second on an M4 Pro MacBook, with peak memory usage of 8.34GiB at an 8K context length. Bailing replicated the core Infinite Wiki experience on a 36GB MacBook Pro, where users click on words while reading to generate contextual explanation cards with multi-layer knowledge drill-down. The demonstration runs entirely on-device with first-token response below 100 milliseconds, and all data remains on the device with no cloud API charges incurred.
On Mac mini, Ling-3.0-tiny is deployed for text-selection translation, grammar correction and text rewriting, with these tasks running offline and delivering low-latency responses. Bailing plans to further strengthen the model's long-tail knowledge coverage and factual accuracy, improve the stability of complex tool calling and long-horizon agent tasks, and enhance its code, scientific reasoning and long-context capabilities. The team also intends to refine hardware adaptation, inference framework support and deployment documentation for the FP8 and INT4 versions, and to build the lightweight model ecosystem together with hardware vendors, universities and the developer community.
Why this event matters
The event has a measured impact on 1 industry. The strongest current signal is positive for Artificial Intelligence, with intensity 60/100 and 70% confidence over a short term horizon.
Impact figures are analytical estimates that combine direction, intensity, confidence and event importance. They are not investment advice.