跳转到主要内容
vllm-project
GitHub 创作者资料

vllm-project

按仓库查看 12 个 GitHub 仓库中的 49 个已收集 skills。

已收集 skills
49
仓库
12
更新
2026-08-29
仓库分布

Skills 分布在哪些仓库

按已收集 skill 数展示主要仓库,并显示它们在该创作者目录中的占比和职业覆盖。

#01
semantic-router
11 个 skills · 2026-08-25
计算机系统分析师市场经理桌面出版专家计算机程序员软件质量保证分析师与测试员

5 个职业分类 · 已分类 100%

22.4%占比
#02
vllm-omni
9 个 skills · 2026-08-29
计算机程序员市场经理综合与运营经理桌面出版专家网页与数字界面设计师

6 个职业分类 · 已分类 100%

18.4%占比
#03
vime
6 个 skills · 2026-06-29
网页与数字界面设计师综合与运营经理

2 个职业分类 · 已分类 100%

12.2%占比
#04
vllm-skills
6 个 skills · 2026-04-03
桌面出版专家计算机程序员

2 个职业分类 · 已分类 100%

12.2%占比
#05
afd-plugin
4 个 skills · 2026-08-26
计算机程序员桌面出版专家市场经理

3 个职业分类 · 已分类 100%

8.2%占比
#06
vllm
3 个 skills · 2026-08-26
网页与数字界面设计师市场经理桌面出版专家

3 个职业分类 · 已分类 100%

6.1%占比
#07
llm-compressor
3 个 skills · 2026-08-27
软件质量保证分析师与测试员综合与运营经理

2 个职业分类 · 已分类 100%

6.1%占比
#08
vllm-ascend
2 个 skills · 2026-08-24
计算机系统分析师其他计算机职业

2 个职业分类 · 已分类 100%

4.1%占比
#09
ci-infra
2 个 skills · 2026-08-05
市场经理综合与运营经理

2 个职业分类 · 已分类 100%

4.1%占比
#10
guidellm
1 个 skills · 2026-08-07
桌面出版专家

1 个职业分类 · 已分类 100%

2%占比
#11
recipes
1 个 skills · 2026-08-26
综合与运营经理

1 个职业分类 · 已分类 100%

2%占比
#12
speculators
1 个 skills · 2026-08-19
市场经理

1 个职业分类 · 已分类 100%

2%占比
仓库浏览

仓库与代表性 skills

openclaw-vsr-bridge
计算机系统分析师

Install vLLM Semantic Router in agent-safe mode, import supported OpenClaw model providers into canonical VSR config, and rewrite OpenClaw to target VSR.

2026-08-18
config-platform-change
市场经理网页与数字界面设计师

Synchronizes config representations across router config, Python CLI schema, and dashboard config UI. Use when adding or changing a config concept that spans those surfaces or addressing config representation debt before Kubernetes-facing translation.

2026-08-03
harness-contract-change
桌面出版专家计算机程序员

Modifies the repository's agent contract including AGENTS.md, docs index, manifests, validation scripts, and contributor-facing harness wrappers. Use when updating agent documentation, changing repo manifests, editing validation scripts, modifying CI/workflow classification, or updating contributor-facing guides like README.md, CONTRIBUTING.md, or the PR template.

2026-08-25
maintainer-issue-pr-management
桌面出版专家

Manages GitHub issue and pull-request lifecycle including creation, updates, triage labelling, and closeout metadata using canonical templates and repository taxonomy. Use when a maintainer asks to create, update, close, or triage GitHub issues or PRs, or when issue creation requires codebase analysis for scope, labels, or acceptance criteria.

2026-08-25
maintainer-release-ops
计算机程序员网页与数字界面设计师

Maintainer release and milestone operating workflow. Use when a maintainer wants to plan a release, assess milestone health, coordinate release blockers, or generate a release-focused review brief.

2026-08-25
routing-calibration-loop
市场经理计算机程序员

Calibrates routing changes against a live router endpoint with executable probes, local DSL validation, versioned deploys, and structured failure review. Use when tuning signals, projections, decisions, or maintained route examples against a real apiserver.

2026-08-03
plugin-end-to-end
软件质量保证分析师与测试员计算机程序员

Implements end-to-end plugin changes spanning router config, post-decision processing, optional CLI/UI exposure, and E2E test coverage. Use when adding a new plugin type, changing plugin config schema or execution semantics, updating plugin chain behavior, or modifying plugin-exposed metadata across surfaces.

2026-08-03
routing-policy-change
计算机程序员

Modifies routing policy after signal extraction, including matched-decision logic, candidate-model selection, and downstream looper behavior. Use when changing decision predicates, thresholds, priorities, model ranking, cost or latency routing, or other post-signal routing policy.

2026-08-03
signal-end-to-end
市场经理

Implements end-to-end signal changes spanning router config, signal extraction, CLI schema, optional bindings, router-owned metadata headers, and E2E test coverage. Use when adding a new signal type, changing signal configuration or extraction logic, updating CLI schema for signal parameters, or modifying router-owned signal metadata contracts.

2026-08-03
startup-chain-change
市场经理

Modifies the local startup chain including image build, container serve/bootstrap logic, and canonical smoke test behavior. Use when changing `vllm-sr serve` behavior, image selection or pull policy, container startup sequences, local Docker/Make bootstrap, or canonical smoke config.

2026-08-03
project-change
桌面出版专家

Handles a focused repository change when no specialized primary skill applies. Use when changed-file routing selects this fallback for a feature, fix, refactor, documentation update, or subsystem-local task.

2026-08-25
add-diffusion-model
计算机程序员

Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including Cache-DiT acceleration and parallelism support (TP, SP/USP, CFG-Parallel, HSDP). Use when integrating a new diffusion model, porting a diffusers pipeline or a custom model repo to vllm-omni, creating a new DiT transformer adapter, adding diffusion model support, or enabling multi-GPU parallelism and cache acceleration for an existing model.

2026-08-29
add-tts-model
市场经理

Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. Use when adding a new TTS model, wiring stage separation for speech synthesis, enabling online voice generation serving, debugging TTS integration behavior, or building audio output pipelines.

2026-08-29
diffusion-perf-opt
综合与运营经理

Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation. Use when Codex is asked to analyze profiling traces, choose parallel strategies, inspect torch profiler trace.json or trace.json.gz timelines, estimate optimization ROI, investigate GPU idle/free bubbles, compare USP/CFG/HSDP/VAE parallelism, or design operator/host/quantization optimizations for vLLM Omni.

2026-05-26
precheck-pr
桌面出版专家

Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness. Use when the user says "precheck", "self review", "pre-submit check", or "check my PR before I open it." Never posts to GitHub.

2026-08-21
quantization
网页与数字界面设计师计算机程序员

Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models. Use when choosing or adding methods such as fp8, int8, gguf, mxfp8, mxfp4, mxfp4_dualscale, ModelOpt, AutoRound, INC, msModelSlim, awq, or gptq; debugging quantized loading; or validating memory, speed, and output quality.

2026-06-10
vllm-omni-npu-model-runner-upgrade
技术写作员

Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.

2026-04-18
vllm-omni-test
桌面出版专家

Generate and run tests for vllm-project/vllm-omni with CI-aligned levels and markers; wire new tests into Buildkite (test-ready.yml for L1/L2, test-merge.yml for L3, test-nightly.yml for L4). On completion, always provide copy-paste local and CI-like pytest commands plus prerequisites. Use when creating regression tests, adding L1-L4 coverage, selecting pytest markers, or validating fixes from issues/PRs.

2026-08-29
review-pr
计算机程序员

Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings. Use for default, detailed, or repeat maintainer reviews; checking correctness, compatibility, tests, benchmarks, model additions, distributed changes, or breaking behavior; and identifying or explicitly requesting the most relevant code-owner reviewers. Use precheck-pr instead for an author's pre-submit self-check.

2026-08-21
find-simplifications
市场经理网页与数字界面设计师

Find evidence-backed simplification candidates in vLLM-Omni and, when requested, turn them into focused proposals or code changes. Use for audits of dead, duplicated, speculative, over-generalized, unnecessarily defensive, or hand-rolled code. Use review-pr for ordinary correctness review and diffusion-perf-opt for performance-first optimization.

2026-08-21
add-dynamic-filter
网页与数字界面设计师

Guide for adding dynamic/filter hooks in vime rollout pipeline. Use when user wants sample-group selection during rollout, buffer filtering before training, or per-sample masking/processing hooks.

2026-06-04
add-eval-dataset-config
网页与数字界面设计师

Guide for adding and validating evaluation dataset configuration in vime. Use when user wants to configure eval datasets via --eval-config or --eval-prompt-data, add per-dataset overrides, or customize evaluation rollout behavior.

2026-06-04
add-reward-function
网页与数字界面设计师

Guide for adding a custom reward function in vime and wiring it through --custom-rm-path (and optional reward post-processing). Use when user wants new reward logic, remote/service reward integration, or task-specific reward shaping.

2026-06-04
add-rollout-function
网页与数字界面设计师

Guide for adding a new rollout function in vime and wiring it through --rollout-function-path. Use when user wants to implement custom rollout data generation logic, custom train/eval rollout outputs, or migrate from the default vLLM rollout path.

2026-06-04
add-tests-and-ci
综合与运营经理

Guide for adding or updating vime tests and CI wiring. Use when tasks require new test cases, CI registration, test matrix updates, or workflow template changes.

2026-06-29
vime-code-review-preferences
综合与运营经理

Use when reviewing or editing vime code, especially refactors around helper APIs, branch selection, argument validation, or recurring reviewer preferences about avoiding unnecessary wrappers and making control flow self-explanatory.

2026-06-29
vllm-bench-random-synthetic
桌面出版专家计算机程序员

Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. Use when the user wants to quickly test vLLM serving performance without downloading external datasets.

2026-04-03
vllm-bench-serve
桌面出版专家

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Supports multiple datasets (random, sharegpt, sonnet, HF), backends (openai, openai-chat, vllm-pooling, embeddings), throughput/latency testing with request-rate control, and result saving. Use when benchmarking LLM serving performance, measuring TTFT/TPOT, or load testing inference APIs.

2026-04-03
vllm-deploy-docker
桌面出版专家

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

2026-04-03
vllm-deploy-k8s
桌面出版专家

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Use this skill whenever the user wants to deploy, run, or serve vLLM on a Kubernetes cluster, including creating deployments, services, checking existing deployments, or managing vLLM on K8s.

2026-04-03
vllm-deploy-simple
计算机程序员

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

2026-04-03
vllm-prefix-cache-bench
桌面出版专家

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Use when the user asks to benchmark prefix caching hit rate, caching efficiency, or repeated-prompt performance in vLLM.

2026-04-03
run-e2e
计算机程序员

Use when the user asks to run, validate, or diagnose the AFD plugin's DeepSeek-V2-Lite GPU/NPU, Qwen3 MoE GPU, or Qwen3.6 MoE CUDA end-to-end tests through the Qwen3.5/3.6 adapter family, including PR-gate E2E, GSM8K-7 accuracy, graph, eager, DBO, or 2A2F scenarios.

2026-08-26
upgrade-gpu-version
桌面出版专家计算机程序员

Upgrade and align the AFD Plugin GPU backend across exact pinned vLLM revisions and CUDA runtime/toolchain environments. Use when Codex must plan, audit, implement, review, or finally validate a GPU vLLM version upgrade; rebase compatibility patches; adapt GPU models, workers, model runners, connectors, CUDA Graph, DBO, TP, profiling, packaging, or isolation contracts; or resolve GPU regressions caused by an upstream version change. Do not use for NPU-only upgrades, model-only adaptation, ordinary GPU bugs unrelated to an upstream upgrade, or E2E-only execution.

2026-08-13
upgrade-npu
市场经理

Upgrade and align the AFD Plugin NPU backend across pinned vLLM and vLLM-Ascend tags or commits. Use when Codex must analyze or implement an NPU runtime upgrade, rebase compatibility patches, adapt NPU models/workers/model runners/connectors, resolve version-driven NPU regressions, or complete the full NPU E2E and accuracy qualification for a new runtime pair. Do not use for GPU-only upgrades, model-only adaptation, ordinary NPU bug fixes unrelated to an upstream upgrade, or E2E-only execution.

2026-08-12
adapt-model
计算机程序员

Guide feasibility analysis, minimal implementation, review, and evidence validation for a new AFD model on GPU, NPU, or both. Use when Codex must adapt or review a new vLLM model, establish its native and AFD execution contract, add model-specific tests, or validate model support. Do not use for ordinary model explanations, existing-only E2E execution, generic bug fixes, or vLLM upgrades.

2026-08-14
RuleHub — Agent Skills 市场与 AI 编程灵感