Get the App
SLTechnology News&Howtos  ›  IT Information  › 

Nvidia announces the new version of TensorRT-LLM: reasoning ability soars 5 times, graphics card above 8GB can run locally, and Chat API of OpenAI is supported.

Shulou Source: shulou.com Published: 2023-11-24 22:03:52 10月03日 Update

CTOnews.com November 16 news, Microsoft Ignite 2023 conference has begun today, Nvidia executives attended the meeting and announced the update of TensorRT-LLM, adding support for OpenAI Chat API.

CTOnews.com reported in October that Nvidia launched Tensor RT-LLM open source libraries for data centers and Windows PC. The biggest feature is that if the Windows PC is equipped with Nvidia GeForce RTX GPU,TensorRT-LLM, the LLM can run four times faster on the Windows PC.

At today's Ignite 2023 conference, Nvidia announced an update to TensorRT-LLM, adding Chat API support for OpenAI, and enhanced DirectML capabilities to improve the performance of AI models such as Llama 2 and Stable Diffusion.

TensorRT-LLM can be done locally through Nvidia's AI Workbench, and developers can use this unified, easy-to-use toolkit to quickly create, test, and customize pre-trained generative AI models and LLM on PC or workstations. Nvidia also launched a pre-emptive experience registration page for this purpose.

Nvidia will release an update to TensorRT-LLM 0.6.0 later this month with a fivefold improvement in reasoning performance and support for other mainstream LLM such as Mistral 7B and Nemotron-3 8B.

Users can run on GeForce RTX 30 series and 40 series GPU with more than 8GB video memory, and some portable Windows devices can also use fast and accurate local LLM functions.

Related readings:

"Nvidia launches Tensor RT-LLM to make large language models run four times faster on PC platforms with RTX."

Tags: Yingwei support run assembly model update function performance speed reasoning maximum to this end mainstream will be in tools toolkits curtains platforms developers data Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno OPPO Reno vpn Redmi Linux Shulou Technology