It was one of the fastest backends last time I checked (with vLLM and lmdeploy being comparable), but the space moves fast. It uses cuda under the hood, torch is not relevant in this context.
https://nvidia.github.io/TensorRT-LLM/performance.html