Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Are there benchmarks on how much faster TensorRT vs native torch/cuda?


https://nvidia.github.io/TensorRT-LLM/performance.html

It was one of the fastest backends last time I checked (with vLLM and lmdeploy being comparable), but the space moves fast. It uses cuda under the hood, torch is not relevant in this context.


I found some official benchmarks for enterprise GPUs, but no comparison data. I couldn't find any benchmarks for commercial GPUs.

https://nvidia.github.io/TensorRT-LLM/performance.html




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: