omlx: Fast LLM Inference
omlx provides fast LLM inference for Apple Silicon projects, enabling continuous batching and SSD caching for efficient machine learning model deployment
- •When building Apple Silicon-based applications that require low-latency LLM inference
- •When developing macOS or iOS apps that rely on large language models
- •When optimizing machine learning workflows for Apple devices
python app.py --model gpt2 --batch-size 16 --ssd-cache /path/to/ssd
Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.
Expect 20–40 minutes in a terminal — or let your AI agent drive it.
mkdir -p ~/.claude/skills/omlx && curl -fsSL https://workflowstacks.com/api/skills/omlx/claude-skill -o ~/.claude/skills/omlx/SKILL.mdOpens the app with this repo with the prompt ready to go — no copy-paste needed.
Omlx: Fast LLM Inference is a very large Python project (~239k lines across 515 code files, plus 403 test files). It is a full software project: use it through its install path rather than reading it end to end. Last commit this month, Apache-2.0 license, has a test suite.
- 1README.mdStart here — what it does and how to install it
- 2omlx/cli.pyWhere the program starts running
- 3pyproject.tomlDependencies and the commands it exposes
- 4omlx/__init__.pyInside omlx/ — the main logic begins here
Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.
Expect 20–40 minutes in a terminal — or let your AI agent drive it.
mkdir -p ~/.claude/skills/omlx && curl -fsSL https://workflowstacks.com/api/skills/omlx/claude-skill -o ~/.claude/skills/omlx/SKILL.mdOpens the app with this repo with the prompt ready to go — no copy-paste needed.
Are you the creator of this tool? Claim your listing → and earn 85% of every sale.
Related skills
More design tools founders pair with this one.