fix: restrict KEDA scaling to /v1/(chat/)?completions paths only

- Switch to HTTPScaledObject (matches a100 pattern)
- Probes (/health, /v1/models) now ignored for scaling
- Keeps fast polling (10s), quick cooldown (30s), 15min idle scale-to-0

Health/readiness/startup probes hit pod IP directly (bypass service), so never triggered scaling anyway — but paths ensure only inference traffic scales.
This commit is contained in:
sirius0xdev 2026-04-29 04:14:46 +00:00
parent d08bf5844a
commit 7e7b7d414d

View file

@ -1,17 +1,22 @@
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
apiVersion: http.keda.sh/v1alpha1
kind: HTTPScaledObject
metadata:
name: rtx6000-scaling
namespace: customer1
spec:
hosts:
- rtx6000-brain-service.customer1.svc.cluster.local
scaleTargetRef:
name: rtx6000-brain-vllm
kind: Deployment
apiVersion: apps/v1
hosts:
- rtx6000-brain-service.customer1.svc.cluster.local
service: rtx6000-brain-service
port: 8000
paths:
- path: "^/v1/(chat/)?completions"
replicas:
min: 0
max: 1
targetPendingRequests: 1
scaledownPeriod: 900\n pollingInterval: 10\n cooldownPeriod: 30
pollingInterval: 10
cooldownPeriod: 30
scaledownPeriod: 900