Skip to main content
Models & Research

China Telecom's Xing4.0 Runs a 29B Coding Agent on 19GB Locally

Subtopic Mixture Of Experts · Long Context · Vision Language Takeaways − China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters. Native 256K context extensible to 512K using MLA attention and 64 routed experts.

By Precis Daily Newsroom1 min read316 words
Illustration for: China Telecom's Xing4.0 Runs a 29B Coding Agent on 19GB
Illustration
Key points
  • Subtopic Mixture Of Experts · Long Context · Vision Language Takeaways − China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters.
  • Native 256K context extensible to 512K using MLA attention and 64 routed experts.
  • First model at this scale trained end-to-end on Huawei Ascend 910C NPUs with MindSpore.

Subtopic Mixture Of Experts · Long Context · Vision Language Takeaways − China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters. Native 256K context extensible to 512K using MLA attention and 64 routed experts. Scores 75.0 on SWE-bench Verified and 57.5 on Terminal-Bench 2.1. First model at this scale trained end-to-end on Huawei Ascend 910C NPUs with MindSpore. GGUF quantizations from 9.9 GB (IQ2_M) to 62.5 GB (F16); Apache 2.0 licensed. See the training report and GitHub repo for details. Xing4.0 brings a 29B mixture-of-experts model to GGUF A GGUF conversion of Xing4.0-29B-A4B has been published on Hugging Face. Developed by China Telecom’s AI team, whose earlier work includes TeleChat, the model contains 29 billion parameters while activating about 4 billion for each token. The release targets local coding agents, repository analysis, and long-context research workloads. Xing4.0 uses a sparse mixture-of-experts architecture with 40 layers and a hidden size of 3,584. Each token is routed through four of 64 specialized experts plus one shared expert. This routing reduces computation per token, while all 29 billion parameters must remain available in system memory, GPU memory, or a combination of both. Multi-head Latent Attention, commonly shortened to MLA, compresses the key-value cache used to track earlier tokens. The model card lists a native context window of 256,000 tokens and extension up to 512,000. MLA makes those windows more practical, although memory use still rises with context length, batch size, and cache precision. The GGUF repository provides several quantizations that trade file size against numerical precision: The model file accounts for only part of runtime memory. Applications also allocate the context cache, compute buffers, model metadata, and backend-specific overhead. CPU offloading can reduce GPU memory requirements, usually with lower generation speed. You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Sources

Summarized from the linked originals.

Related stories