gcloud-lab/infrastructure/gpus/base
sirius0xdev 7e7b7d414d fix: restrict KEDA scaling to /v1/(chat/)?completions paths only
- Switch to HTTPScaledObject (matches a100 pattern)
- Probes (/health, /v1/models) now ignored for scaling
- Keeps fast polling (10s), quick cooldown (30s), 15min idle scale-to-0

Health/readiness/startup probes hit pod IP directly (bypass service), so never triggered scaling anyway — but paths ensure only inference traffic scales.
2026-04-29 04:14:46 +00:00
..
keda-gpu-scaling fix: restrict KEDA scaling to /v1/(chat/)?completions paths only 2026-04-29 04:14:46 +00:00
vllm-servers feat: optimize RTX6000 vLLM KEDA scaling 2026-04-29 04:12:36 +00:00