Add 12 AI infrastructure and LLM operations skills

New skills covering hot-topic AI engineering subjects:

Local AI Infrastructure:
- vllm-server: High-throughput LLM serving with PagedAttention, tensor parallelism, quantization
- llm-inference-scaling: KEDA-based GPU autoscaling for LLM inference on Kubernetes
- rag-infrastructure: Production RAG with hybrid search, reranking, and embedding pipelines
- llm-fine-tuning: QLoRA/LoRA fine-tuning with Axolotl, DeepSpeed ZeRO-3, and DPO alignment

Infrastructure:
- gpu-server-management: NVIDIA driver setup, MIG partitioning, DCGM monitoring
- vector-database-ops: Qdrant, Weaviate, pgvector for production AI search
- llm-gateway: LiteLLM-based API gateway with rate limiting, virtual keys, fallback routing

DevOps/AI:
- llm-cost-optimization: Model right-sizing, prompt/semantic caching, batch API, break-even analysis
- llm-caching: Multi-layer exact + semantic + provider caching to cut costs 30-70%
- ai-pipeline-orchestration: Prefect/Airflow/Dagster for RAG ingestion and training workflows

Orchestration:
- model-serving-kubernetes: KServe + Triton with canary deployments and GPU autoscaling

Security:
- ai-security-hardening: Prompt injection defense, PII scrubbing, model weight verification

https://claude.ai/code/session_011MN1C4PrkCeg2Qmi7q1ZUe
This commit is contained in:
Claude
2026-03-02 01:18:29 +00:00
parent cc3848d963
commit dd77232b16
13 changed files with 3327 additions and 0 deletions
+12
View File
@@ -195,6 +195,7 @@ No agent? No problem. Browse the skills, copy the scripts, use the configs. It's
| [argocd-gitops](devops/orchestration/argocd-gitops/) | GitOps with ArgoCD |
| [kustomize](devops/orchestration/kustomize/) | Kubernetes manifest customization |
| [openshift](devops/orchestration/openshift/) | OpenShift cluster management |
| [model-serving-kubernetes](devops/orchestration/model-serving-kubernetes/) | KServe and Triton model serving with canary deployments and GPU autoscaling |
### Observability
| Skill | Description |
@@ -212,6 +213,9 @@ No agent? No problem. Browse the skills, copy the scripts, use the configs. It's
|-------|-------------|
| [agent-observability](devops/ai/agent-observability/) | Tracing, latency, token, and cost telemetry for agents |
| [agent-evals](devops/ai/agent-evals/) | Automated regression and safety eval suites for agents |
| [llm-cost-optimization](devops/ai/llm-cost-optimization/) | Cut LLM API costs with caching, batching, model routing, and self-hosting |
| [llm-caching](devops/ai/llm-caching/) | Exact and semantic caching layers to reduce API calls by 3070% |
| [ai-pipeline-orchestration](devops/ai/ai-pipeline-orchestration/) | Orchestrate RAG ingestion, training, and batch inference with Prefect/Airflow |
### Release Management
| Skill | Description |
@@ -276,6 +280,7 @@ No agent? No problem. Browse the skills, copy the scripts, use the configs. It's
|-------|-------------|
| [ai-agent-security](security/ai/ai-agent-security/) | Defend agents against injection, tool abuse, and exfiltration |
| [llm-app-security](security/ai/llm-app-security/) | Harden LLM app inputs, outputs, and tenant isolation |
| [ai-security-hardening](security/ai/ai-security-hardening/) | Harden LLM deployments against prompt injection, model theft, and data exfiltration |
</details>
@@ -334,6 +339,7 @@ No agent? No problem. Browse the skills, copy the scripts, use the configs. It's
| [user-management](infrastructure/servers/user-management/) | Users, groups, sudo |
| [systemd-services](infrastructure/servers/systemd-services/) | Services and timers |
| [performance-tuning](infrastructure/servers/performance-tuning/) | System optimization |
| [gpu-server-management](infrastructure/servers/gpu-server-management/) | NVIDIA GPU driver setup, MIG partitioning, DCGM monitoring for AI workloads |
### Networking
| Skill | Description |
@@ -343,6 +349,7 @@ No agent? No problem. Browse the skills, copy the scripts, use the configs. It's
| [cdn-setup](infrastructure/networking/cdn-setup/) | CloudFront, Cloudflare |
| [reverse-proxy](infrastructure/networking/reverse-proxy/) | nginx, Traefik |
| [service-mesh](infrastructure/networking/service-mesh/) | Istio, Linkerd |
| [llm-gateway](infrastructure/networking/llm-gateway/) | Unified LLM API gateway with routing, rate limiting, virtual keys, and semantic caching |
### Databases
| Skill | Description |
@@ -353,6 +360,7 @@ No agent? No problem. Browse the skills, copy the scripts, use the configs. It's
| [mongodb](infrastructure/databases/mongodb/) | MongoDB clusters |
| [redis](infrastructure/databases/redis/) | Redis caching |
| [database-backups](infrastructure/databases/database-backups/) | Backup strategies |
| [vector-database-ops](infrastructure/databases/vector-database-ops/) | Qdrant, Weaviate, and pgvector for production AI search and RAG workloads |
### Storage
| Skill | Description |
@@ -375,6 +383,10 @@ No agent? No problem. Browse the skills, copy the scripts, use the configs. It's
| [ollama-stack](infrastructure/local-ai/ollama-stack/) | Private local inference stack with Ollama |
| [mac-mini-llm-lab](infrastructure/local-ai/mac-mini-llm-lab/) | Mac mini setup for always-on local LLM serving |
| [openclaw-local-mac-mini](infrastructure/local-ai/openclaw-local-mac-mini/) | OpenClaw setup for local development and Mac mini hosting |
| [vllm-server](infrastructure/local-ai/vllm-server/) | High-throughput LLM serving with vLLM — PagedAttention, tensor parallelism, OpenAI API |
| [llm-inference-scaling](infrastructure/local-ai/llm-inference-scaling/) | Auto-scale LLM inference clusters on Kubernetes with KEDA and GPU-aware scheduling |
| [rag-infrastructure](infrastructure/local-ai/rag-infrastructure/) | Production RAG with vector stores, hybrid search, embedding pipelines, and reranking |
| [llm-fine-tuning](infrastructure/local-ai/llm-fine-tuning/) | QLoRA and full fine-tuning with Axolotl, DeepSpeed, and DPO alignment on GPU clusters |
### IT Operations
| Skill | Description |