Overview
- Jefferies, citing Silicon Data, said Monday that average inference costs fell to about US$1.16–US$1.20 per million tokens, the lowest level recorded this year.
- A surge in Chinese open‑weight and open‑source models has lowered operating costs by letting firms download and self‑host models rather than pay closed APIs.
- Major U.S. labs have cut prices on specific models to defend share, with OpenAI reducing developer rates for its GPT‑5.6 Luna model by roughly 80 percent.
- Some Chinese providers are using mixed strategies by offering cheap open models while raising fees for certain user groups such as programmers, showing market segmentation.
- The price shift is speeding adoption of local and agentic AI uses, squeezing vendor margins and contributing to investor reevaluations of large cloud AI players while geopolitical access rules could reshape competition.