Skip to content
Back to insights

Technical Case Study · Backend Engineering

Go in production: architecture, caching, and operations

Architecture, caching, and observability choices in Hospital Sírio-Libanês’ Go backend, which handles over 20 million requests per month.

The backend connects bedside TVs, the Tasy ERP, and streaming infrastructure at the São Paulo and Brasília sites. This article covers decisions to limit request cost, handle integration failures, and observe the service in production. The 6 ms average response refers to the API, not the complete TV navigation or playback experience.

Observed production scale

requests per month
20M+
average API response
6 ms
sustained cache hit rate
92%
backend engineering commits
1k+

System view

An architecture with clear degradation paths

Timeouts and retry limits bound the time spent on integrations. The cache combines Redis and local memory to reduce dependence on a single layer during failures.

HTTP Edge (Fiber v2)
Go Core API
Hybrid Cache (Redis + sync.Map)
PostgreSQL (pgxpool/sqlc)
External Platforms (TASY / IPTV)
Prometheus & Dashboard
In this article

01 · Context

Real challenges of a hospital ecosystem

An unstable bedside TV creates support work and affects the patient experience. The backend needed to integrate Tasy and IPTV streaming while keeping resource use and failure behavior predictable.

  • Bidirectional integration with TASY ERP via HSL ESB for patient bed activation, deactivation, and device swap.
  • IPTV streaming middleware orchestration with dynamic salted MD5 authentication tokens.
  • Continuous Android TV navigation backed by aggressive MAC address caching.
  • Internal admin panel secured by HttpOnly JWT cookies with comprehensive audit logging.

02 · Architecture

Predictable bootstrap and strict boundaries

We chose Fiber v2 and FastHTTP with a focus on throughput and memory use. Application configuration is established at startup, before the service accepts traffic.

  • Fiber v2 and FastHTTP with buffer reuse to reduce unnecessary allocations.
  • High-concurrency PostgreSQL access via pgx/v5 and pgxpool, using sqlc for core type-safe queries and GORM for admin.
  • Explicit timeouts at every layer (HTTP Read/Write, DB connection lifetime, and dial timeouts).
  • Readiness probes validating PostgreSQL and Redis health before routing traffic to new instances.
Failing fast under extreme load is far superior to accumulating goroutines until an OOM container kill occurs.

03 · Performance

Less work on the main request path

The route serving the catalogue and bed data is the focus of optimization. We reduced memory allocations, reused connections, and tracked latency to evaluate the result.

  • Custom high-performance HTTP client (FastHTTP) with per-host connection pools and a 3-retry limit.
  • JSON processing with Sonic JSON and typed queries generated by sqlc, reducing repetitive work on the request path.
  • Selective compression and ETag headers to conserve internal hospital network bandwidth.
  • Monitoring dashboard and HTML templates compiled directly into the Go binary via embed.FS.

04 · Caching

Two-Tier Hybrid Caching (Redis + sync.Map)

Redis is the primary cache layer. To handle outages or slow responses, the application also uses local memory, with expiration and cleanup rules.

  • Primary Redis (v8) pool with volatility-driven TTLs (e.g., 4 min for MAC login cache).
  • Automatic fallback to local sync.Map with a TTL janitor routine when Redis fails or exceeds 500ms.
  • Dedicated background goroutine executing daily preventive purges between 03:00 and 04:00 AM.
  • Real-time hit/miss/error telemetry broken down by tier (Redis vs Local).
Fallback reduces dependence on Redis, but memory belongs to each instance. Data validity and the capacity of this layer also need to be considered when evaluating failures.

05 · Observability

Normalized metrics and embedded Live Dashboard

Full telemetry exported natively for Prometheus and visualized in a real-time system dashboard served directly by the Go binary.

  • http_request_duration_seconds histograms with custom latency buckets (0.5ms to 30s) and normalized routes.
  • Dedicated metrics for database queries (db_query_duration_seconds) and external integrations (ESB/IPTV).
  • Real-time infrastructure metrics: CPU utilization, heap/stack memory, goroutines, and active pool connections.
  • Responsive web dashboard served at /pkg/dashboard embedded via embed.FS with zero external dependencies.
Mandatory route parameter normalization in middleware prevented metric cardinality explosion in Prometheus.

06 · Security

4-Tier RBAC Matrix and Audit Trail

JWTMiddleware and AuthorizeMiddleware enforce access control. Roles define permitted actions, and audit records support investigations into changes.

  • Admin session cookies (admin_token) protected with HttpOnly, Secure, SameSite, and signed via golang-jwt/jwt/v5.
  • 4-tier RBAC matrix (Dev, Support, Manager, Analyst) with technical guardrails preventing managers from escalating privileges.
  • PostgreSQL audit log (activity_log) recording actor ID, action type (CREATE/UPDATE/DELETE), target entity, and field deltas.
  • Employee passwords hashed with Bcrypt (golang.org/x/crypto/bcrypt) and secret rotation policies.

07 · Incidents

Practical lessons learned in production

Production incidents informed changes to request limits, caching, and instrumentation. These are the issues and responses that shaped the service.

  • Prometheus metric cardinality explosion → Solution: strict endpoint normalization middleware.
  • Redis network blips → Solution: transparent hybrid cache with automatic sync.Map local fallback.
  • Unbounded ESB retries → Solution: encapsulated FastHTTP client with exponential backoff and finite attempts.
  • Inconsistencies when replacing a TV → response: check the MAC address before persisting the device in the database.

08 · Checklist

What to verify before the next deployment

Before release, we review operational limits, simulate failures, and check behavior under load. Average latency alone does not describe every service condition.

  • pgxpool and Redis connection limits tuned against available CPU container cores.
  • Cache fallback mechanism verified under fault injection (simulated Redis outage).
  • Prometheus metrics audited for normalized routes with zero dynamic ID leakage in labels.
  • Multi-stage Docker build compiled with -ldflags='-s -w' and load-tested prior to release.

Architecture has to work beyond the diagram

Architecture becomes clearer when decisions and outcomes appear together.

Continue through the case studies to see how the same criteria appear across other production contexts.

Explore workRoberto Moraes · Software Engineer