China Telecom's Apache 2.0 MoE, trained on Huawei Ascend, scores 75.00 on SWE-bench Verified in its own table, a point behind Qwen3.6, and its 4-bit file is ~18 GB.
China Telecom's AI subsidiary has released Xing4.0-29B-A4B, an open-weight mixture-of-experts model that it pitches as a coding and agent model you can run on one consumer graphics card. The weights went up on Hugging Face on 16 September 2026 under the Apache 2.0 licence, and the formal announcement followed on 24 September. The headline claim is a score of 75.00 on SWE-bench Verified, with only 4 billion of the model's 29 billion parameters active for each token.
Two things in the vendor's own material temper that headline. In the model card's comparison table, Alibaba's Qwen3.6-35B-A3B scores 76.00 on the same benchmark, so Xing4.0 is second, not first. And the "15 GB of GPU memory" figure in the press release does not match the official 4-bit file, which the card describes as about 18 GB. Every score below is the vendor's own. As of 2 October we found no independent reproduction.
What was released
Xing4.0 is the next generation of the series previously called TeleChat, built by China Telecom Artificial Intelligence Technology Co., Ltd. The model card lists:
| Spec | Xing4.0-29B-A4B |
|---|---|
| Parameters | 29B total, 4B active per token |
| Layers | 40 |
| Attention | MLA (multi-head latent attention) |
| Experts | 64 routed, 4 active per token, plus 1 shared |
| Context | 256K native, extensible to 512K |
| Licence | Apache 2.0 |
The card says the architecture combines mHC, MLA and MTP (multi-token prediction). Besides the BF16 weights, China Telecom published an FP8 checkpoint and an IQ4_NL GGUF build. The card lists Transformers, vLLM, SGLang and KTransformers for inference, and LLaMA-Factory and MindFormers for fine-tuning. It also says the chat format was adapted for agent tools including OpenCode, Claude Code, OpenClaw and Hermes.
The model is popular already. The main repository shows 1.83k likes and 49,408 downloads in the last month, and community GGUF, MLX and abliterated (safety-removed) variants appeared within days.
Trained without Nvidia
The more significant claim may be about how the model was made, not how it scores. China Telecom says Xing4.0 is "the first model of this scale trained entirely on the Ascend NPU platform with the MindSpore framework". That means Huawei's Ascend 910C clusters, not Nvidia GPUs. The card says a set of optimisations, among them MoE communication tuning, selective recomputation and fused operators, raised training throughput by about 96% over the out-of-the-box setup.
The card offers no training-cost or cluster-size figures to check that against. Still, a competitive 29B MoE trained on domestic accelerators is a data point for anyone following how far export controls constrain Chinese model development.
From Ascend training to local inference. The official 4-bit GGUF is about 18 GB, more than the 15 GB in the press release. Sources: Xing4.0 model cards, China Telecom AI press release.
The benchmark table, read carefully
China Telecom compares Xing4.0 with two similar-sized open MoE models: Google's Gemma4-26B-A4B and Alibaba's Qwen3.6-35B-A3B. All three columns come from the same model-card table:
| Benchmark | Xing4.0-29B-A4B | Gemma4-26B-A4B | Qwen3.6-35B-A3B |
|---|---|---|---|
| SWE-bench Verified | 75.00 | 53.00 | 76.00 |
| SWE-bench Multilingual | 66.00 | 51.00 | 67.20 |
| Terminal-Bench 2.1 | 57.50 | 30.00 | 51.50 |
| Claw-Eval | 76.55 | 71.49 | 74.54 |
| DeepresearchBII | 60.80 | 39.30 | 59.70 |
| Tau3-Bench | 64.63 | 58.90 | 67.20 |
| AIME2026 | 90.00 | 88.30 | 92.70 |
| IFBench | 69.67 | 72.67 | 65.50 |
| AA.LCR | 61.00 | 66.00 | 62.00 |
Read across the rows, the picture is narrower than "the new number to beat":
- Where Xing4.0 leads clearly: Terminal-Bench 2.1 (57.50 against Qwen's 51.50) and, by smaller margins, Claw-Eval and DeepresearchBII. These are the agentic, tool-using tests the model was tuned for.
- Where it trails Qwen3.6: SWE-bench Verified by one point, SWE-bench Multilingual by 1.2, plus Tau3-Bench and AIME2026.
- Against Gemma4-26B-A4B, the closest match on active parameters, it is far ahead on every coding and agent benchmark. It is behind only on IFBench and AA.LCR.
Xing4.0-29B-A4B against Qwen3.6-35B-A3B, benchmark by benchmark. All scores are vendor-reported in the Xing4.0 model card.
The footnotes say how Xing4.0 was tested: SWE-bench Verified in the SWE-agent harness at temperature 1.0 with a 210K context window, and Terminal-Bench 2.1 in terminus-2 averaged over three runs. They do not say whether Gemma4 and Qwen3.6 were run by China Telecom under the same settings or copied from those vendors' own reports. SWE-bench scores move several points with the scaffold, so treat the one-point gap to Qwen as a tie, not a ranking.
The 15 GB claim
The press release says the model "requires only 15 GB of GPU memory" and can "run locally on a single consumer-grade graphics card", using "low-bit quantization and memory optimization techniques". It does not name a quantization format or a context length.
The official GGUF repository holds a single IQ4_NL file that its card calls "approximately 18 GB" (Hugging Face's file list shows 20.1 GB). With the KV cache on top, that will not fit inside a 16 GB card. A 24 GB card is the realistic floor for running it fully on the GPU at that quantization. On a 16 GB card, expect to offload some expert layers to system RAM with llama.cpp or KTransformers. Because only 4B parameters are active per token, that costs less speed than it would for a dense 29B model, but it is not the "fits in 15 GB" experience the release implies. Smaller community quantizations may get under 16 GB, but the vendor's benchmark scores were not measured on them.
What to do with it
- For local coding agents: Xing4.0 is worth a trial next to Qwen3.6-35B-A3B, especially for terminal-heavy agent work, where its Terminal-Bench lead is the largest gap in the table. Run both on your own repositories in the harness you actually use. Vendor SWE-bench numbers do not carry across scaffolds.
- Size the hardware to the real file. Plan on a 24 GB card for the official IQ4_NL build, or accept partial CPU offload on 16 GB.
- Check provenance before you deploy. Pull from
XingChen-AGI, the official organisation. Dozens of re-uploads, fine-tunes and abliterated variants are already on Hugging Face, and an abliterated build has had its refusal behaviour removed on purpose. - Treat it like any open-weight model in an agent. A coding agent that can run shell commands needs the same sandboxing, egress limits and secret hygiene whichever model sits behind it.
- Wait for independent numbers before you replace a model you have already validated. A one-point gap on a vendor-run benchmark is not a reason to switch.