In 2026, the world will reach a significant milestone where AI compute used for inference and running models surpasses the compute required for training. This shift is driven by the increasing demand for agents to perform multi-step reasoning, tool usage, and data access, which requires extensive GPU resources.