Skill · em Dados, IA e pesquisa
serving-llms-vllm
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
Procedência
- Origem: davila7/claude-code-templates
- Caminho:
cli-tool/components/skills/ai-research/inference-serving-vllm - Versão fixada:
57f899e5394bb8ca166f38eacae8f0853cbfe033 - Licença: MIT
- Espelhado em 25/09/2026
- nenhum download no Claude Code Templates (lido em 25/09/2026)
Antes de instalar
5 arquivos · 35 KB · só texto, nenhum script
Instalar na sua CLI
O comando baixa a versão fixada (commit 57f899e) direto da origem, para a pasta que a CLI lê. Precisa de curl (macOS e Linux); no Windows não há comando, porque o Rook Labs é para macOS.
Claude Code
Neste projeto: instala em .claude/skills/serving-llms-vllm/.
d=".claude/skills/serving-llms-vllm" u="https://raw.githubusercontent.com/davila7/claude-code-templates/57f899e5394bb8ca166f38eacae8f0853cbfe033/cli-tool/components/skills/ai-research/inference-serving-vllm" curl -fsSL --create-dirs \ -o "$d/SKILL.md" "$u/SKILL.md" \ -o "$d/references/optimization.md" "$u/references/optimization.md" \ -o "$d/references/quantization.md" "$u/references/quantization.md" \ -o "$d/references/server-deployment.md" "$u/references/server-deployment.md" \ -o "$d/references/troubleshooting.md" "$u/references/troubleshooting.md"
Global: instala em ~/.claude/skills/serving-llms-vllm/.
d="$HOME/.claude/skills/serving-llms-vllm" u="https://raw.githubusercontent.com/davila7/claude-code-templates/57f899e5394bb8ca166f38eacae8f0853cbfe033/cli-tool/components/skills/ai-research/inference-serving-vllm" curl -fsSL --create-dirs \ -o "$d/SKILL.md" "$u/SKILL.md" \ -o "$d/references/optimization.md" "$u/references/optimization.md" \ -o "$d/references/quantization.md" "$u/references/quantization.md" \ -o "$d/references/server-deployment.md" "$u/references/server-deployment.md" \ -o "$d/references/troubleshooting.md" "$u/references/troubleshooting.md"
Codex
Neste projeto: instala em .agents/skills/serving-llms-vllm/.
d=".agents/skills/serving-llms-vllm" u="https://raw.githubusercontent.com/davila7/claude-code-templates/57f899e5394bb8ca166f38eacae8f0853cbfe033/cli-tool/components/skills/ai-research/inference-serving-vllm" curl -fsSL --create-dirs \ -o "$d/SKILL.md" "$u/SKILL.md" \ -o "$d/references/optimization.md" "$u/references/optimization.md" \ -o "$d/references/quantization.md" "$u/references/quantization.md" \ -o "$d/references/server-deployment.md" "$u/references/server-deployment.md" \ -o "$d/references/troubleshooting.md" "$u/references/troubleshooting.md"
Global: instala em ~/.agents/skills/serving-llms-vllm/.
d="$HOME/.agents/skills/serving-llms-vllm" u="https://raw.githubusercontent.com/davila7/claude-code-templates/57f899e5394bb8ca166f38eacae8f0853cbfe033/cli-tool/components/skills/ai-research/inference-serving-vllm" curl -fsSL --create-dirs \ -o "$d/SKILL.md" "$u/SKILL.md" \ -o "$d/references/optimization.md" "$u/references/optimization.md" \ -o "$d/references/quantization.md" "$u/references/quantization.md" \ -o "$d/references/server-deployment.md" "$u/references/server-deployment.md" \ -o "$d/references/troubleshooting.md" "$u/references/troubleshooting.md"
Antigravity
Neste projeto: instala em .agents/skills/serving-llms-vllm/.
d=".agents/skills/serving-llms-vllm" u="https://raw.githubusercontent.com/davila7/claude-code-templates/57f899e5394bb8ca166f38eacae8f0853cbfe033/cli-tool/components/skills/ai-research/inference-serving-vllm" curl -fsSL --create-dirs \ -o "$d/SKILL.md" "$u/SKILL.md" \ -o "$d/references/optimization.md" "$u/references/optimization.md" \ -o "$d/references/quantization.md" "$u/references/quantization.md" \ -o "$d/references/server-deployment.md" "$u/references/server-deployment.md" \ -o "$d/references/troubleshooting.md" "$u/references/troubleshooting.md"
Global: instala em ~/.gemini/antigravity-cli/skills/serving-llms-vllm/.
d="$HOME/.gemini/antigravity-cli/skills/serving-llms-vllm" u="https://raw.githubusercontent.com/davila7/claude-code-templates/57f899e5394bb8ca166f38eacae8f0853cbfe033/cli-tool/components/skills/ai-research/inference-serving-vllm" curl -fsSL --create-dirs \ -o "$d/SKILL.md" "$u/SKILL.md" \ -o "$d/references/optimization.md" "$u/references/optimization.md" \ -o "$d/references/quantization.md" "$u/references/quantization.md" \ -o "$d/references/server-deployment.md" "$u/references/server-deployment.md" \ -o "$d/references/troubleshooting.md" "$u/references/troubleshooting.md"
Peça ao Rook
Já usa o Rook Labs? Cole no chat do Rook: instale a skill https://rooklabs.sh/marketplace/cct.serving-llms-vllm
Prévia do SKILL.md
---
name: serving-llms-vllm
description: Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
version: 1.0.0
author: Orchestra Research
license: MIT
tags: [vLLM, Inference Serving, PagedAttention, Continuous Batching, High Throughput, Production, OpenAI API, Quantization, Tensor Parallelism]
dependencies: [vllm, torch, transformers]
---
# vLLM - High-Performance LLM Serving
## Quick start
vLLM achieves 24x higher throughput than standard transformers through PagedAttention (block-based KV cache) and continuous batching (mixing prefill/decode requests).
**Installation**:
```bash
pip install vllm
```
**Basic offline inference**:
```python
from vllm import LLM, SamplingParams
llm = LLM(model="meta-llama/Llama-3-8B-Instruct")
sampling = SamplingParams(temperature=0.7, max_tokens=256)
outputs = llm.generate(["Explain quantum computing"], sampling)
print(outputs[0].outputs[0].text)
```
**OpenAI-compatible server**:
```bash
vllm serve meta-llama/Llama-3-8B-Instruct
# Query with OpenAI SDK
python -c "
from openai import OpenAI
client = OpenAI(base_url='http://localhost:8000/v1', api_key='EMPTY')
…