According to a new industry survey by renowned analyst Ming-Chi Kuo, NVIDIA has restarted the "Rubin CPX" AI inference prefill acceleration GPU project and significantly adjusted its product design, targeting the high-cost prefill stage of AI inference. The new Rubin CPX's computing power is close to that of the standard Rubin GPU, with a maximum power consumption of 2300 watts per card and a switch to 168GB HBM4 memory. The project is expected to begin production in Q1 2027, aiming to reduce long-context inference costs through dedicated hardware and heterogeneous deployment. Kuo states that over 50% of current AI inference workloads come from prefill and KV Cache construction.