ai-agent

TensorRT-Edge-LLM: Fast AI Inference

Get high-performance AI inference with TensorRT-Edge-LLM, for founders building physical AI products.
484 stars94 forksPythonGuide quality 8/10Updated 7/23/2026100% free · open source
What it does

TensorRT-Edge-LLM enables high-performance AI inference for physical AI products by optimizing large language models and vision models for edge devices

When to use it
  • Building a smart home device that requires real-time voice or image processing
  • Developing an autonomous robot that needs to quickly process visual or auditory data
  • Creating a wearable device that requires low-latency AI-driven insights
Ready-to-paste prompt
Run the sample inference application with a custom prompt: ./sample_inference -m models/bert-base-uncased.pt -p 'What is the capital of France?'
Heads up: Ensure you have the necessary NVIDIA hardware and software dependencies installed, including CUDA and cuDNN, to run TensorRT-Edge-LLM
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

✅ Light setup

Installs with a command or two; your AI agent can do it for you.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/tensorrt-edge-llm && curl -fsSL https://workflowstacks.com/api/skills/tensorrt-edge-llm/claude-skill -o ~/.claude/skills/tensorrt-edge-llm/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

Quick Actions
Details
Creator
NVIDIA
Language
Python
Category
ai-agent
Published
10/2/2025

Are you the creator of this tool? Claim your listing → and earn 85% of every sale.