gcloud-lab/README.md
Hermes Agent 6183159974 feat: add Prometheus, Grafana, and Tailscale monitoring stack
- Install Prometheus + Grafana via kube-prometheus-stack (ClusterIP only, no public ingress)
- Deploy Tailscale Operator for secure VPN access to internal services
- Add CNPG/PostgreSQL monitoring dashboards
- Add vLLM inference monitoring dashboards (tokens, latency, GPU)
- Add Cilium networking dashboards (policy, traffic, drops)
- Update infra-controllers staging kustomization to include all controllers
- Add monitoring-configs Flux sync for dashboard deployment
- Update README with monitoring architecture and access instructions
- Remove broken stale monitoring files (Azure Key Vault refs, wrong domains)

Access: kubectl port-forward or Tailscale VPN (replace auth key before deploy)
2026-04-26 02:31:53 +00:00

16 KiB

GCloud-Lab DevOps Infrastructure

A cloud-native DevOps laboratory project showcasing modern infrastructure-as-code, GitOps practices, and Kubernetes orchestration on Google Cloud Platform. This project runs a news intelligence system with LLM-powered analysis and a workflow automation platform.

Table of Contents


Project Overview

This repository contains infrastructure and application configurations for:

  1. News Intelligence Pipeline: Automated web scraping, LLM-powered summarization, and Telegram distribution
  2. Workflow Automation: N8N platform for custom integrations
  3. DevOps Reference Architecture: Demonstrates GitOps, IaC, and cloud-native best practices

Architecture

┌─────────────────────────────────────────────────────────────────────────┐
│                        Google Cloud Platform                            │
│  ┌───────────────────────────────────────────────────────────────────┐  │
│  │                    GKE Cluster (devops-lab-cluster)               │  │
│  │                                                                   │  │
│  │  ┌─────────────┐  ┌─────────────┐  ┌─────────────┐               │  │
│  │  │ Standard    │  │ L4 GPU Pool │  │ A100 GPU    │               │  │
│  │  │ Node Pool   │  │ (SPOT L4)   │  │ (SPOT A100) │               │  │
│  │  │ e2-std-2    │  │ 1 node (24/7│  │ 0-1 nodes   │               │  │
│  │  └─────────────┘  └─────────────┘  └─────────────┘               │  │
│  │                                                                   │  │
│  │  ┌─────────────────────────────────────────────────────────────┐ │  │
│  │  │                    Cilium CNI + Hubble                      │ │  │
│  │  └─────────────────────────────────────────────────────────────┘ │  │
│  │                                                                   │  │
│  │  ┌─────────────────────────────────────────────────────────────┐ │  │
│  │  │                   customer1 namespace                       │ │  │
│  │  │  - OpenClaw PAaaS (Dual-Tier vLLM: L4 Dispatch / A100 Think)│ │  │
│  │  │  - News Bot Pipeline (DeepSeek-R1 Quant Analyst)            │ │  │
│  │  │  - PAaaS Landing Page (GHCR Docker Pulls)                   │ │  │
│  │  │  - CloudNativePG Isolated Databases                         │ │  │
│  │  └─────────────────────────────────────────────────────────────┘ │  │
│  └───────────────────────────────────────────────────────────────────┘  │
└─────────────────────────────────────────────────────────────────────────┘

DevOps Tools & Technologies

Infrastructure as Code (IaC)

Tool Version Purpose
Terraform 1.7+ Infrastructure provisioning for GCP resources
Google Provider 7.14.1 Terraform provider for GCP
Helm Provider Latest Terraform provider for Helm charts
Flux Provider 1.7.6 Terraform provider for Flux bootstrap

Container Orchestration & Networking

Tool Version Purpose
Google Kubernetes Engine (GKE) Latest Managed Kubernetes cluster
Cilium 1.18.5 CNI plugin with eBPF-based networking
Hubble 1.18.5 Network observability and monitoring
Kubernetes Gateway API v1 Ingress routing and traffic management

GitOps & Configuration Management

Tool Version Purpose
Flux CD 1.7.6 GitOps continuous delivery
Kustomize v1beta1 Kubernetes manifest customization
Helm 3+ Kubernetes package manager
SOPS Latest Secrets encryption in Git
Age Latest Modern encryption for SOPS

Database

Tool Version Purpose
CloudNative PG 0.26.1 PostgreSQL Kubernetes operator
PostgreSQL 15.2 Relational database (3-node HA cluster)

AI/ML Infrastructure

Tool Version Purpose
Ollama Latest Local LLM inference server
Gemma2 Latest Open-source LLM for text summarization
NVIDIA L4 GPU - GPU acceleration for LLM workloads

Development Environment

Tool Version Purpose
Mise Latest Development tool version manager
Dev Containers Latest Consistent development environment
k9s Latest Kubernetes CLI dashboard

Monitoring & Observability

Tool Version Purpose
Prometheus Latest Metrics collection via kube-prometheus-stack
Grafana Latest Dashboards & visualizations
Tailscale Latest Secure VPN access to internal services

Project Structure

gcloud-lab/
├── modules/                          # Terraform IaC modules
│   ├── providers.tf                  # Provider configurations
│   ├── gke.tf                        # GKE cluster definition
│   ├── vpc.tf                        # VPC and subnet configuration
│   ├── nodepool.tf                   # Standard node pool
│   ├── nodepool-gpu.tf               # GPU node pool (SPOT instances)
│   ├── flux.tf                       # Flux GitOps bootstrap
│   ├── helm.tf                       # Helm chart deployments (Cilium)
│   └── variables.tf                  # Input variables
│
├── clusters/                         # Cluster configurations
│   └── devops-lab/
│       ├── flux-system/              # Flux CD components
│       │   ├── gotk-components.yaml  # Flux controllers
│       │   ├── gotk-sync.yaml        # Git repository sync
│       │   └── kustomization.yaml    # Flux kustomization
│       ├── customer1.yaml            # Customer1 Kustomization
│       ├── infra-controllers.yaml    # Infrastructure controllers (CNPG, KEDA, Monitoring, Tailscale)
│       └── infra-configs.yaml        # Infrastructure configs
│
├── infrastructure/                   # Infrastructure components
│   ├── controllers/
│   │   ├── base/
│   │   │   ├── cnpg/                 # CloudNative PG operator
│   │   │   ├── keda/                 # KEDA autoscaling
│   │   │   ├── monitoring/           # Prometheus + Grafana (no public ingress)
│   │   │   └── tailscale/            # Tailscale Operator for secure VPN access
│   │   └── staging/
│   │       └── kustomization.yaml    # Aggregates all base components
│   └── configs/
│       └── staging/
│           └── kustomization.yaml
│
├── apps/                             # Application deployments
│   ├── base/
│   │   └── customer1/
│   │       ├── namespace.yaml        # Namespace definition
│   │       ├── deployment.yaml       # N8N deployment
│   │       ├── service.yaml          # ClusterIP service
│   │       ├── storage.yaml          # PersistentVolumeClaim
│   │       ├── configmap.yaml        # N8N configuration
│   │       ├── pg-cluster-customer1.yaml  # PostgreSQL cluster
│   │       ├── apigateway.yaml       # GCP Gateway
│   │       ├── http-route.yaml       # HTTP routing
│   │       ├── healthcheck.yaml      # Health check policy
│   │       └── news_bot/             # News bot microservices
│   │           ├── scraper-cronjob.yaml
│   │           ├── analyst-cronjob.yaml
│   │           ├── telebot-cronjob.yaml
│   │           ├── scrapy-configmap.yaml
│   │           └── scrapy-urls-configmap.yaml
│   └── staging/
│       └── customer1/
│           └── kustomization.yaml    # Staging overlay
│
├── scripts/
│   └── setup                         # Development setup script
│
├── .devcontainer.json                # Dev container configuration
├── mise.toml                         # Tool version management
├── age.agekey                        # SOPS encryption key
└── README.md                         # This file

Infrastructure Components

GKE Cluster

  • Name: devops-lab-cluster
  • Region: us-central1-a
  • Network: Custom VPC with dual-stack IPv4/IPv6

Node Pools

Pool Machine Type Scaling Purpose
Standard e2-standard-2 1-16 nodes General workloads
GPU (SPOT) g2-standard-8 + L4 0-5 nodes LLM inference

Networking

  • VPC: devops-lab-network
  • Primary CIDR: 10.0.0.0/16
  • Pod CIDR: 192.168.32.0/20
  • Service CIDR: 192.168.16.0/24
  • CNI: Cilium with advanced datapath
  • Ingress: GCP L7 Global Load Balancer

GitOps Flow

GitHub Repository
       │
       ▼
  Flux Source Controller (watches git, 1min interval)
       │
       ▼
  Flux Kustomize Controller (applies manifests)
       │
       ├── infrastructure/controllers → CNPG, KEDA, Monitoring, Tailscale
       ├── infrastructure/configs     → Cluster configs
       └── apps/staging/customer1     → Applications

Applications

1. Private Assistant as a Service (PAaaS)

A premium, uncensored, privacy-first AI assistant platform with dual-tier cognitive architecture:

  • Tier 1 (Dispatcher): L4 GPU Spot instance running 24/7. Hosts Qwen2.5-Coder-7B-Instruct-heretic via vLLM v0.9.1 for lightning-fast, cheap triage and tool calling (using the pythonic tool parser).
  • Tier 2 (Deep Thinker): A100 80GB Spot instance scaling from 0-1 via KEDA. Hosts Qwen3.5-27B-heretic with --enable-chunked-prefill and --kv-cache-dtype=fp8 for massive multi-file context and reasoning without OOMing or stalling concurrent users.
  • Frontend: Isolated openclaw deployments per tenant, connected to Telegram/Discord via outbound polling (no public ingress required).
  • Landing Page: Dockerized marketing site built via CI/CD from openclaw-projects and deployed to the staging kustomization overlay.

2. Autonomous News Quant Pipeline (news_bot)

An institutional-grade pipeline scraping 371 global feeds to generate actionable futures trading signals:

  • Scraper: CronJob at :50 pulling multi-lingual global financial data.
  • Map/Reduce Analyst: Utilizes DeepSeek-R1 (with a strict 10-step <think> protocol) and local open-weights to extract "Market-Moving DNA". Translates events into explicit futures targets (/ES, /CL, /NQ) with R:R, Take Profit, and Stop Loss levels anchored in provided volume/price data.
  • Privacy: All proprietary technical data stays strictly within the VPC, executing against local models rather than public APIs like OpenAI to protect the trading edge and avoid throttling during market panics.

3. N8N Workflow Automation

  • Database: PostgreSQL (dedicated n8n database)
  • Custom integrations and webhook catchers.

(Note: PineScript trading strategies have been migrated out of this IaC repository and live in openclaw-projects/trading-bots.)

Getting Started

Prerequisites

  • Google Cloud account with billing enabled
  • GitHub account with repository access
  • gcloud CLI authenticated
  • Terraform 1.7+

Local Development Setup

# Install tools via mise
./scripts/setup

# Or manually
mise trust && mise install

Infrastructure Deployment

cd modules

# Initialize Terraform
terraform init

# Set required variables
export TF_VAR_github_token="your-token"
export TF_VAR_github_org="your-org"
export TF_VAR_github_repository="gcloud-lab"

# Plan and apply
terraform plan
terraform apply

Accessing the Cluster

# Configure kubectl
gcloud container clusters get-credentials devops-lab-cluster \
  --zone us-central1-a \
  --project devops-lab-cluster

# Verify connection
kubectl get nodes

# Use k9s for interactive management
k9s

Accessing Monitoring (Grafana / Prometheus)

Monitoring services are not publicly exposed. Access is via Tailscale VPN or port-forwarding:

# Option 1: Port-forward Grafana
kubectl port-forward svc/prometheus-community-kube-prometheus-stack-grafana \
  -n monitoring 3000:3000

# Option 2: Port-forward Prometheus
kubectl port-forward svc/prometheus-community-kube-prometheus-stack-prometheus \
  -n monitoring 9090:9090

⚠️ Before deploying, replace the Grafana admin password in infrastructure/controllers/base/monitoring/release.yaml with a secure value, or create a monitoring-grafana-admin Secret instead.


Security

Secrets Management

  • Encryption: SOPS with Age encryption
  • Key Storage: age.agekey (do not commit unencrypted)
  • Flux Integration: Automatic decryption during deployment

Pod Security

  • Non-root containers (UID 1000)
  • Filesystem group enforcement
  • Privilege escalation disabled
  • Resource limits enforced

Network Security

  • Cilium network policies for pod-to-pod isolation
  • TLS termination at load balancer
  • Private cluster networking with NAT

Database Security

  • Managed roles with secret-based passwords
  • Separate users per application (customer1, news_app)
  • HA cluster with automatic failover

Tool Reference

Terraform Providers

google      = "~> 7.14"   # GCP resources
helm        = "~> 2.0"    # Helm chart management
flux        = "~> 1.7"    # GitOps bootstrap

Helm Charts

cilium:           1.18.5    # CNI and service mesh
cloudnative-pg:   0.26.1    # PostgreSQL operator

Container Images

docker.n8n.io/n8nio/n8n:2.1.4
ghcr.io/cloudnative-pg/postgresql:15.2
ollama/ollama:latest
siriussec/newsscraper:latest
siriussec/summarizer:latest
siriussec/news-messenger:latest

Cost Optimization

  • SPOT GPU Instances: 60-90% savings on LLM workloads
  • Autoscaling: GPU nodes scale to 0 when idle
  • Resource Limits: Prevents runaway costs
  • Scheduled Workloads: CronJobs only run when needed

License

Private repository - All rights reserved.