Cutting YOLOv8 inference latency down through TensorRT compilation, INT8 calibration, and a container layout that avoids cold-start penalties.