Nvidia news
On July 13, 2026, technical benchmarks were released for the GLM--5.2 large language model running on a 4× GB10 cluster within the new DGX Spark architecture. The system demonstrated a decode speed of 22 tokens per second with a 256K context window, highlighting significant improvements in long--context processing efficiency. This performance milestone underscores NVIDIA's continued leadership in providing optimized hardware for next--generation agentic AI and high--throughput enterprise workloads.