sirius0xdev
b40ffc68cb
update release
2026-05-04 01:17:40 +00:00
sirius0xdev
dc400a364d
fix tsproxy issue
2026-05-04 01:09:48 +00:00
sirius0xdev
896ecc8c51
Merge branch 'master' of github.com:sirius0xdev/gcloud-lab
...
i am the captain #
2026-05-04 01:07:18 +00:00
sirius0xdev
3874014c85
fix kustomizations
2026-05-04 01:07:09 +00:00
sirius0xdev
bf4a2b260f
fix: increase tailscale-operator HelmRelease timeout to 15m
2026-05-04 00:50:43 +00:00
sirius0xdev
2b6225c4e8
fix: increase tailscale-operator HelmRelease timeout to 10m (install was timing out)
2026-05-04 00:48:09 +00:00
sirius0xdev
852d2c76c1
fix: update tailscale operator authkey with fresh single-use key
2026-05-04 00:43:59 +00:00
sirius0xdev
61165c2b00
fix: separate monitoring from controllers so broken prometheus-stack doesn't block tailnet chain
2026-05-04 00:30:02 +00:00
sirius0xdev
83b70951c6
fix(tailnet): move TsProxy to separate Kustomization that depends on operator
...
- Remove TsProxy from infrastructure-controllers to avoid CRD timing issue
- Create infrastructure/tailnet/ Kustomization for TsProxy resources
- Add infrastructure-tailnet Flux Kustomization with dependsOn
- Add dependsOn to customer1 Kustomization for trade-dashboard TsProxy
2026-05-04 00:06:51 +00:00
sirius0xdev
5f9b55af55
feat(tailscale): add operator authkey secret and rtx6000-brain TsProxy
...
- Add SOPS-encrypted tailscale-operator-authkey secret for operator auth
- Add TsProxy to expose rtx6000-brain-service on tailnet (port 8000)
- Enable trade-dashboard TsProxy (was waiting for operator install)
2026-05-03 23:30:28 +00:00
sirius0xdev
bd760287a8
fix(rtx6000): enable scale-to-zero with 20min cooldown + wake-up route
...
- Set replicas.min: 0 for scale-to-zero when idle
- Add config.cooldownPeriod: 1200 to KEDA HTTP add-on HelmRelease
(20 min buffer before scaling down from 1 replica)
- Add external HTTPRoute at brain.siriusdevops.com for manual wake-up
via curl from Hermes/Telegram
2026-04-29 16:01:03 +00:00
Hermes Agent
6183159974
feat: add Prometheus, Grafana, and Tailscale monitoring stack
...
- Install Prometheus + Grafana via kube-prometheus-stack (ClusterIP only, no public ingress)
- Deploy Tailscale Operator for secure VPN access to internal services
- Add CNPG/PostgreSQL monitoring dashboards
- Add vLLM inference monitoring dashboards (tokens, latency, GPU)
- Add Cilium networking dashboards (policy, traffic, drops)
- Update infra-controllers staging kustomization to include all controllers
- Add monitoring-configs Flux sync for dashboard deployment
- Update README with monitoring architecture and access instructions
- Remove broken stale monitoring files (Azure Key Vault refs, wrong domains)
Access: kubectl port-forward or Tailscale VPN (replace auth key before deploy)
2026-04-26 02:31:53 +00:00
sirius0xdev
7a3126caff
working on certmanager and promethus stack
2026-04-23 01:19:56 +00:00
Sirius Claw
4fb80d3157
fix: update KEDA Helm controllers to Flux v2 APIs
2026-04-20 04:00:10 +00:00
Sirius Claw
e736d5f874
feat: Add KEDA and KEDA HTTP Add-on controllers
2026-04-20 03:36:34 +00:00
sirius0xdev
b6d93886bf
fix infra kustomization.yaml
2026-02-02 23:28:51 +00:00
sirius0xdev
995baa64cb
added service account for db backups and changed terraform to include bucket and service accounts
2026-02-02 23:20:46 +00:00
sirius0xdev
b0eb3c2e8c
fix release.yaml
2026-01-06 01:54:18 +00:00
sirius0xdev
87b2dae39f
fix cnpg kustomization
2026-01-06 01:46:55 +00:00
sirius0xdev
521ea092af
fix dir structure
2026-01-06 01:30:10 +00:00