ai-agent

Run Large LLMs with RTX 6000 Pro

Run large LLMs on PCIe GPUs with RTX 6000 Pro, a tool for founders working with AI agents.
905 stars63 forksPythonGuide quality 9/10Updated 8/31/2026100% free · open source
What it does

RTX 6000 Pro allows founders to run large language models like Qwen3.5-397B, Kimi-K2.5, and GLM-5 on PCIe GPUs without NVLink, enabling local inference for AI agents

When to use it
  • You need to deploy large language models on local hardware for data privacy or latency reasons
  • Your AI workflow requires running models like Qwen3.5-397B or Kimi-K2.5 on commodity GPUs
  • You want to test and fine-tune large language models without relying on cloud services
Ready-to-paste prompt
python run_qwen.py --model qwen3.5-397b --gpu 0 --batch-size 32
Heads up: Ensure you have a compatible NVIDIA driver installed (version 470 or later) and that your system meets the minimum hardware requirements outlined in the RTX 6000 Pro Wiki
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

✅ Light setup

Installs with a command or two; your AI agent can do it for you.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/rtx6kpro && curl -fsSL https://workflowstacks.com/api/skills/rtx6kpro/claude-skill -o ~/.claude/skills/rtx6kpro/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

Quick Actions
Details
Creator
local-inference-lab
Language
Python
Category
ai-agent
Published
3/9/2026

Are you the creator of this tool? Claim your listing → and earn 85% of every sale.