ai-agent

Kimi Audio

Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
4,720 stars375 forksPythonUpdated 6/21/2025100% free · open source
What it does

Kimi-Audio is an open-source audio foundation model that enables startup founders to build applications with audio understanding, generation, and conversation capabilities.

When to use it
  • When building a voice assistant or chatbot that needs to understand audio inputs
  • When creating a podcast or audio content platform that requires automatic transcription or summarization
  • When developing an application that generates music or audio effects
Ready-to-paste prompt
python infer.py --input_audio 'path/to/input/audio.wav' --output_text 'path/to/output/text.txt'
Heads up: The model requires a significant amount of computational resources and memory to train, so ensure you have a suitable GPU and sufficient RAM before attempting to train the model
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

✅ Light setup

Installs with a command or two; your AI agent can do it for you.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/kimi-audio && curl -fsSL https://workflowstacks.com/api/skills/kimi-audio/claude-skill -o ~/.claude/skills/kimi-audio/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

Quick Actions
Details
Creator
MoonshotAI
Language
Python
Category
ai-agent
Published
4/25/2025

Are you the creator of this tool? Claim your listing → and earn 85% of every sale.