← All work

Platform & infrastructure · Observability

Running blockchain infrastructure at 99.99%

Zenith Chain · Abuja · 2020–2021

Focus
Platform & infrastructure · Observability
Evidence
4 documented proof points
Decision record
Firsthand trade-off included

The brief

What was built

Nodes, wallet services and indexers on Kubernetes across 20+ microservices, with event-driven autoscaling absorbing high-throughput transaction spikes without degrading response times, and reusable Terraform modules behind every environment.

Evidence

Recorded outcomes

  • 99.99% uptime across blockchain nodes, wallet services and indexers — high-availability configuration and automated failover, on checks that measure consensus progress rather than process health
  • Cloud spend down 35% through KEDA event-driven autoscaling and right-sizing compute against real-time workload demand instead of provisioning for peak
  • Deployment frequency up 3× with zero-downtime releases across production blockchain and fintech workloads
  • Incident detection and diagnosis cut 30% on Prometheus, Grafana, Loki and OpenTelemetry; environment provisioning cut 60% through reusable Terraform modules

Decision record

Healthy is not the same as caught up

99.99% cost more than redundant nodes. One RPC node stayed “healthy” while falling 1,800 blocks behind — the process was up, so every check passed. I replaced process-up alerts with block-height, peer-count and consensus-progress checks. The same pass went the other way: after a harmless 3am state-pruning spike, I stopped paging on CPU alone.

Return to all work