Microsoft Azure and University of Texas researchers found that multi-step AI workflows create CPU-GPU bottlenecks that conventional inference infrastructure struggles to handle efficiently.