gcloud-lab/infrastructure/gpus/base/vllm-servers
sirius0xdev d08bf5844a feat: optimize RTX6000 vLLM KEDA scaling
- Rename deployment to rtx6000-brain-vllm to match scaler target ref
- Set deployment replicas to 0 (KEDA controlled)
- Add pollingInterval: 10s and cooldownPeriod: 30s for faster response

This enables scale-to-zero when idle and quick scale-up on first request.

See https://github.com/sirius0xdev/gcloud-lab/tree/master/infrastructure%2Fgpus%2Fbase
2026-04-29 04:12:36 +00:00
..
a100-vllm.yaml clean up and add routes to paas site 2026-04-28 03:31:52 +00:00
kustomization.yaml fix typos 2026-04-28 03:33:52 +00:00
rtx6000-vllm.yaml feat: optimize RTX6000 vLLM KEDA scaling 2026-04-29 04:12:36 +00:00
vllm-gemma.yaml clean up and add routes to paas site 2026-04-28 03:31:52 +00:00
vllm-l4.yaml clean up and add routes to paas site 2026-04-28 03:31:52 +00:00