ai-agent

SALMONN

SALMONN family: A suite of advanced multi-modal LLMs
1,526 stars124 forksUpdated 8/24/2026100% free · open source
What it does

SALMONN is a family of multi‑modal large language models that can understand and generate text, images, and video in a single request.

When to use it
  • You need an open‑source model that can answer questions about an image or video (e.g., “What product is shown in this photo?”).
  • You want to prototype a visual‑assistant feature without paying for a commercial API.
  • Your startup requires on‑premise inference for privacy‑sensitive visual data.
Ready-to-paste prompt
python inference.py --image_path examples/product.jpg --prompt "What are the key features of this product and how would you market it to millennials?"
Heads up: The model requires CUDA 11.8+ and a GPU with at least 24 GB VRAM; attempting to run on a CPU will raise an out‑of‑memory error.
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

✅ Light setup

Installs with a command or two; your AI agent can do it for you.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/salmonn && curl -fsSL https://workflowstacks.com/api/skills/salmonn/claude-skill -o ~/.claude/skills/salmonn/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

Quick Actions
Details
Creator
bytedance
Category
ai-agent
Published
8/11/2023

Are you the creator of this tool? Claim your listing → and earn 85% of every sale.