SparkLLM and iFLYTEK Unveil Longest Context On-Device Models with Million-Token Capacity

September 1, 2026
SparkLLM and iFLYTEK Unveil Longest Context On-Device Models with Million-Token Capacity
  • The models support end-to-end local operation on diverse hardware (NVIDIA, Huawei, Hygon, Houmo) and integrate with popular deployment ecosystems, enabling offline use and incremental training.

  • Training relied on hundreds of billions of tokens of long-document data, enabling the models to maintain context across multi-step interactions without requiring cloud connectivity.

  • Beyond long context, these models offer edge-native capabilities for personal office work, code development, smart devices, and robotics, including agent tooling and tool-calling features.

  • A hybrid attention architecture optimized for agents, code, mathematics, and instruction following underpins these models, designed to handle long documents and specialized material.

  • Upcoming: Spark X2.5-293B is set to release on September 7, with further upgrades to code and agent capabilities.

  • SparkLLM has released and open-sourced two on-device general models, Spark X2.5-4B and Spark X2.5-1.7B, each boasting a native context window of up to 1,000,000 tokens—the longest context yet for on-device models.

  • iFLYTEK’s subsidiary Ciyuan Xinghuo open-sourced two edge-side general LLMs, Xinghuo X2.5-4B and Xinghuo X2.5-1.7B, with native 1,000,000-token context, and made model weights, code, and deployment docs fully public.

  • Early demonstrations show capabilities across tasks like analyzing sales spreadsheets with Loomy, generating bilingual reports, code-related work with local tooling, high-accuracy domestic smart-home controls, and robot/edge applications with reduced cloud reliance.

  • Both models run on multiple hardware platforms—including NVIDIA, Huawei, Haiguang, and Houmo—and are compatible with vLLM, SGLang, and llama.cpp, with quick deployment options via Ollama and LM Studio, and incremental training through LLaMA-Factory.

  • Training involved about 20 trillion tokens on a fully domestic compute platform, using high-quality supervised fine-tuning and reinforcement learning, with availability through open-source and commercial channels.

  • Upcoming: iFLYTEK will release Xinghuo X2.5, its next-generation flagship LLM, on September 7 to advance code and agent capabilities.

  • Performance highlights include near cloud-scale coding capabilities for X2.5-4B and high end-to-end control accuracy on Domux smart-home tasks (about 90.3% accuracy, 0.85 seconds average response).

Summary based on 2 sources


Get a daily email with more Tech stories

More Stories