Svmuu News: OpenAI’s engineering team recently revealed to some colleagues that the company has found a new method to optimize its systems, which can reduce the “inference” cost of AI models by more than half. Inference cost refers to the computational resources consumed by a model when it is actually running and responding to user requests. This optimization stems primarily from improved efficiency in utilizing existing server resources, rather than relying on investments in new computing chips. This development reflects how AI companies, while continuing to compete for computing resources, are also enhancing the efficiency of their existing infrastructure through software and system-level optimizations to alleviate the pressure of rapidly rising model operating costs. (The Information)