gcloud-lab/infrastructure/gpus
sirius0xdev 7e7b7d414d fix: restrict KEDA scaling to /v1/(chat/)?completions paths only
- Switch to HTTPScaledObject (matches a100 pattern)
- Probes (/health, /v1/models) now ignored for scaling
- Keeps fast polling (10s), quick cooldown (30s), 15min idle scale-to-0

Health/readiness/startup probes hit pod IP directly (bypass service), so never triggered scaling anyway — but paths ensure only inference traffic scales.
2026-04-29 04:14:46 +00:00
..
base fix: restrict KEDA scaling to /v1/(chat/)?completions paths only 2026-04-29 04:14:46 +00:00
staging Update kustomization.yaml 2026-04-28 23:51:19 -04:00