China Telecom's Xing4.0 Runs a 29B Coding Agent on 19GB Locally
Subtopic Mixture Of Experts · Long Context · Vision Language China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters. Native 256K context extensible to 512K using MLA attention and 64 routed experts.

- Subtopic Mixture Of Experts · Long Context · Vision Language China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters.
- Native 256K context extensible to 512K using MLA attention and 64 routed experts.
- First model at this scale trained end-to-end on Huawei Ascend 910C NPUs with MindSpore.
Subtopic Mixture Of Experts · Long Context · Vision Language
China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters.
Native 256K context extensible to 512K using MLA attention and 64 routed experts.
Scores 75.0 on SWE-bench Verified and 57.5 on Terminal-Bench 2.1.
First model at this scale trained end-to-end on Huawei Ascend 910C NPUs with MindSpore.
GGUF quantizations from 9.9 GB (IQ2_M) to 62.5 GB (F16); Apache 2.0 licensed.
See the training report and GitHub repo for details.
Xing4.0 brings a 29B mixture-of-experts model to GGUF
A GGUF conversion of Xing4.0-29B-A4B has been published on Hugging Face. Developed by China Telecom’s AI team, whose earlier work includes TeleChat, the model contains 29 billion parameters while activating about 4 billion for each token. The release targets local coding agents, repository analysis, and long-context research workloads.
Xing4.0 uses a sparse mixture-of-experts architecture with 40 layers and a hidden size of 3,584. Each token is routed through four of 64 specialized experts plus one shared expert. This routing reduces computation per token, while all 29 billion parameters must remain available in system memory, GPU memory, or a combination of both.
Multi-head Latent Attention, commonly shortened to MLA, compresses the key-value cache used to track earlier tokens. The model card lists a native context window of 256,000 tokens and extension up to 512,000. MLA makes those windows more practical, although memory use still rises with context length, batch size, and cache precision.
The GGUF repository provides several quantizations that trade file size against numerical precision:
The model file accounts for only part of runtime memory. Applications also allocate the context cache, compute buffers, model metadata, and backend-specific overhead. CPU offloading can reduce GPU memory requirements, usually with lower generation speed.
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
Sources
Related stories

Components of A Coding Agent
How Coding Agents Use Tools, Memory, and Repo Context to Make LLMs Work Better in Practice Sebastian Raschka, PhD Apr 04, 2026 978 65 100 In this article, I want to cover the overall design of coding agents and agent harnesses: what they are, how they work, and how the different.

Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Model by Half With FP8
Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.

Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI
Search agents powered by large language models (LLMs) are transforming how enterprises retrieve information.

SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation
- Apple Machine Learning Research research area Methods and Algorithms , research area Tools, Platforms, Frameworks content type paper published September 2026 SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation Authors Youngmok Jung, Sirajul Salekin, Henry.

How Botika runs full-stack generative AI on Modal
Botika builds agentic e-commerce teams that automate visual production for global fashion brands. Powered by custom models researched and trained in-house, Botika handles the entire pipeline - from 4K image generation to real-time personalization at scale.

Vulnerabilities in AI Agents Expose Trust Gap Flaw
Google and other organizations have acknowledged vulnerabilities in their AI agents, which exploit trust gaps in the Model Context Protocol.

Cohere's North Small Translate Beats DeepL Across 50 Languages With Downloadable Weights
Subtopic Dpo · Fine Tuning · Distillation Cohere released North Small Translate , a 218B / 25B-active MoE translation model with open weights. Scores 83.60 on WMT26 across 50+ languages, beating DeepL NextGen (81.37) and Google Translate (68.20).

Add secure Web Search to Claude Desktop with Amazon Bedrock AgentCore
Claude Desktop on Amazon Bedrock provides powerful AI assistance, but without integrated web search, responses are limited to the model’s training knowledge cutoff. When you need current information, such as recent documentation updates, live pricing, or weather updates, the.