Awesome Go

pgwd

CategoryDatabase
SubcategoryDatabase Tools
Stars7

CLI that monitors PostgreSQL connection counts (total, active, idle, stale) and notifies via Slack and/or Loki when thresholds are exceeded. Supports Kubernetes (kubectl port-forward) and optional run context in notifications

About pgwd

This large README is shown as a plain-text preview. Read the full README at the source

# pgwd — Postgres Watch Dog

<a id="top"></a>

<p align="center">
  <img src="docs/logo.svg" alt="pgwd" width="96" height="96">
</p>

<p align="center">
  <em>Watch your PostgreSQL connections</em>
</p>

[![Version](https://img.shields.io/badge/version-1.2.0-blue)](https://github.com/hrodrig/pgwd/releases)
[![Release](https://img.shields.io/github/v/release/hrodrig/pgwd)](https://github.com/hrodrig/pgwd/releases)
[![CI](https://github.com/hrodrig/pgwd/actions/workflows/ci.yml/badge.svg)](https://github.com/hrodrig/pgwd/actions)
[![codecov](https://codecov.io/gh/hrodrig/pgwd/graph/badge.svg)](https://codecov.io/gh/hrodrig/pgwd)
[![gghstats clones](https://gghstats.hermesrodriguez.com/api/v1/badge/hrodrig/pgwd?metric=clones)](https://gghstats.hermesrodriguez.com/hrodrig/pgwd)
[![Mentioned in Awesome Go](https://awesome.re/mentioned-badge.svg)](https://github.com/avelino/awesome-go)
[![Go 1.26.6](https://img.shields.io/badge/go-1.26.6-00ADD8?logo=go)](https://go.dev/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![pkg.go.dev](https://pkg.go.dev/badge/github.com/hrodrig/pgwd)](https://pkg.go.dev/github.com/hrodrig/pgwd)
[![deps.dev](https://img.shields.io/badge/deps.dev-go%20module-blue)](https://deps.dev/go/github.com%2Fhrodrig%2Fpgwd)
[![DEV.to](https://img.shields.io/badge/DEV.to-Article-0A0A0A?logo=dev.to)](https://dev.to/hrodrig/pgwd-a-watchdog-for-your-postgresql-connections-1pjg)
[![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/hrodrig/pgwd)

**Repo:** [github.com/hrodrig/pgwd](https://github.com/hrodrig/pgwd) · **Releases:** [GitHub Releases](https://github.com/hrodrig/pgwd/releases) · **Spec:** [SPECIFICATIONS.md](SPECIFICATIONS.md) · **Operator:** [pgwd-selfhosted](https://github.com/hrodrig/pgwd-selfhosted) · **Changelog:** [CHANGELOG.md](CHANGELOG.md) · **Roadmap:** [ROADMAP.md](ROADMAP.md) · **Article:** [pgwd on DEV — a watchdog for your PostgreSQL connections](https://dev.to/hrodrig/pgwd-a-watchdog-for-your-postgresql-connections-1pjg)

**Supply chain (from v0.8.0):** each release attaches **SPDX** and **CycloneDX** SBOMs plus **Cosign** signatures for **`checksums.txt`** and **`ghcr.io/hrodrig/pgwd`** images — see [Supply chain verification](#supply-chain-verification).

Go CLI that checks PostgreSQL connection counts (active/idle) and notifies via **Slack** and/or **Loki** when configured thresholds are exceeded. It can also alert on **stale connections** (connections that stay open and never close).

### Why connection limits matter

PostgreSQL enforces a configured ceiling (`max_connections`). Slots reserved for privileged roles (`superuser_reserved_connections`, and on PostgreSQL 16+ `reserved_connections`) reduce how many ordinary application connections can succeed before the server starts refusing new sessions. When the limit is hit, clients fail to connect with SQLSTATE **53300** (“too many clients”). Size connection pools and reserved slots so applications fail loudly before total exhaustion rather than silently degrading. Under high concurrency, each backend is still a process: memory (`shared_buffers` plus per-session/`work_mem` pressure), CPU, and I/O can degrade even before hard rejection—or the OS can OOM.

**pgwd** watches connection pressure (and connect failures) so you can alert before or when saturation happens. It does **not** replace connection poolers or careful sizing. For hardware/app heuristics, reserved slots, WAL/standby notes, and developer pool pitfalls, see **[PostgreSQL connection limits and saturation](docs/postgresql-connection-limits.md)**.

**Self-hosted deployment (Docker Compose, Helm, Kubernetes manifests):** **[pgwd-selfhosted](https://github.com/hrodrig/pgwd-selfhosted)** — production paths, env layout, and observability stacks live there; this repo ships the application binary, packages, and container image only.

**Related tools (same maintainer):**
- **[pgwd](https://github.com/hrodrig/pgwd)** — PostgreSQL connection watchdog ([live traffic](https://gghstats.hermesrodriguez.com/hrodrig/pgwd); deploy: [pgwd-selfhosted](https://github.com/hrodrig/pgwd-selfhosted))
- **[gghstats](https://github.com/hrodrig/gghstats)** — GitHub repo traffic beyond 14 days ([live demo](https://gghstats.hermesrodriguez.com); deploy: [gghstats-selfhosted](https://github.com/hrodrig/gghstats-selfhosted))
- **[kzero](https://github.com/hrodrig/kzero)** — bastion-first declarative workload reset ([live traffic](https://gghstats.hermesrodriguez.com/hrodrig/kzero); deploy: [kzero-selfhosted](https://github.com/hrodrig/kzero-selfhosted))
- **[groot](https://github.com/hrodrig/groot)** — Kubernetes diagnostics archive ([live traffic](https://gghstats.hermesrodriguez.com/hrodrig/groot); deploy: [groot-selfhosted](https://github.com/hrodrig/groot-selfhosted))

**Documentation:** [ROADMAP.md](ROADMAP.md), [Operator use cases](docs/use-cases.md), [Connection limits](docs/postgresql-connection-limits.md), [SPECIFICATIONS.md](SPECIFICATIONS.md) (behavior contract), [Kubernetes passwords / DISCOVER migration](docs/kubernetes-passwords.md), [docs/](docs/README.md) (band plans, sequence diagrams, upgrades), `man pgwd`. **Scanning:** [tools/README.md](tools/README.md).

![Terminal demo](docs/demo.gif)

### How pgwd fits together

Click the diagram for the **interactive** viewer (guided views, pan/zoom, Present mode, source links, flow animation). GitHub cannot run the HTML inline — the preview below is a static snapshot.

<p align="center">
  <a href="https://htmlpreview.github.io/?https://raw.githubusercontent.com/hrodrig/pgwd/develop/docs/pgwd-architecture.html">
    <img src="docs/pgwd-architecture.png" alt="pgwd architecture — config, check loop, notifiers, store, Kubernetes port-forward" width="900">
  </a>
</p>

<p align="center">
  <a href="https://htmlpreview.github.io/?https://raw.githubusercontent.com/hrodrig/pgwd/develop/docs/pgwd-architecture.html"><strong>Open interactive architecture →</strong></a>
  · or clone and open <code>docs/pgwd-architecture.html</code> locally
</p>

Config → check loop → PostgreSQL · notifiers → Slack/Loki/PagerDuty/Teams/webhook · store → `/metrics` + CSV · optional Kubernetes port-forward. Spec: [`docs/architecture.pgwd.json`](docs/architecture.pgwd.json).

## Table of contents

- [How pgwd fits together](#how-pgwd-fits-together)
- [Quick start](#quick-start)
- [Compare](#compare)
- [Configuration: CLI vs environment](#configuration-cli-vs-environment)
- [Usage examples](#usage-examples)
- [Typical scenarios](#typical-scenarios)
- [Kubernetes](#kubernetes)
- [Parameters](#parameters)
- [Install](#install)
- [Build](#build)
- [Testing](#testing)
- [Requirements](#requirements)
- [Slack](#slack)
- [Loki](#loki)
- [Troubleshooting](#troubleshooting)
- [FAQ](#faq)
- [Docker](#docker)
- [systemd](#systemd)
- [Debian / Ubuntu](#debian--ubuntu)
- [AlmaLinux](#almalinux)
- [OpenSUSE](#opensuse)
- [Arch Linux](#arch-linux)
- [Alpine Linux (OpenRC)](#alpine-linux-openrc)
- [OpenBSD](#openbsd)
- [FreeBSD](#freebsd)
- [NetBSD](#netbsd)
- [DragonFly BSD](#dragonfly-bsd)
- [Solaris](#solaris)
- [Roadmap](#roadmap)
- [Get involved](#get-involved)
- [License](#license)

---

## Quick start

```bash
# See all options
pgwd -h

# Minimal: check once, alert to Slack (3-tier levels 75/85/95% by default)
pgwd -db-url "postgres://user:pass@localhost:5432/mydb" \
     -notifications-slack-webhook "https://hooks.slack.com/services/..."

# Custom 3-tier levels (default 75,85,95)
pgwd -db-url "postgres://..." -notifications-slack-webhook "https://..." -db-threshold-levels 70,85,90
```

### Breaking changes (upgrade from 0.9.x)

**1.0.0** removes legacy `db:`, total/active thresholds, and `notify-on-connect-failure`. Migration checklist: **[docs/UPGRADE-0.9-to-1.0.md](docs/UPGRADE-0.9-to-1.0.md)**.

### Breaking changes (upgrade from 0.5.x)

If you use CLI flags or env vars for notifications or DB thresholds, update your scripts:

| Old | New |
|-----|-----|
| `-threshold-total` | removed; use `-db-threshold-levels` |
| `-threshold-active` | removed; use `-db-threshold-levels` |
| `-threshold-idle` | `-db-threshold-idle` |
| `-threshold-stale` | `-db-threshold-stale` |
| `-threshold-levels` | `-db-threshold-levels` |
| `-stale-age` | `-db-stale-age` |
| `-default-threshold-percent` | `-db-default-threshold-percent` |
| `-slack-webhook` | `-notifications-slack-webhook` |
| `-loki-url` | `-notifications-loki-url` |
| `-loki-labels` | `-notifications-loki-labels` |
| `-loki-org-id` | `-notifications-loki-org-id` |
| `-loki-bearer-token` | `-notifications-loki-bearer-token` |
| `PGWD_SLACK_WEBHOOK` | `PGWD_NOTIFICATIONS_SLACK_WEBHOOK` |
| `PGWD_LOKI_URL` | `PGWD_NOTIFICATIONS_LOKI_URL` |
| `PGWD_LOKI_LABELS` | `PGWD_NOTIFICATIONS_LOKI_LABELS` |
| `PGWD_LOKI_ORG_ID` | `PGWD_NOTIFICATIONS_LOKI_ORG_ID` |
| `PGWD_LOKI_BEARER_TOKEN` | `PGWD_NOTIFICATIONS_LOKI_BEARER_TOKEN` |
| `PGWD_THRESHOLD_TOTAL` | removed; use `PGWD_DB_THRESHOLD_LEVELS` |
| `PGWD_THRESHOLD_ACTIVE` | removed; use `PGWD_DB_THRESHOLD_LEVELS` |
| `PGWD_THRESHOLD_IDLE` | `PGWD_DB_THRESHOLD_IDLE` |
| `PGWD_THRESHOLD_STALE` | `PGWD_DB_THRESHOLD_STALE` |
| `PGWD_THRESHOLD_LEVELS` | `PGWD_DB_THRESHOLD_LEVELS` |
| `PGWD_STALE_AGE` | `PGWD_DB_STALE_AGE` |
| `PGWD_DEFAULT_THRESHOLD_PERCENT` | `PGWD_DB_DEFAULT_THRESHOLD_PERCENT` |

Config file keys unchanged.

**Consolidated upgrade guide (0.5.x → 0.6.x):** [docs/UPGRADE-0.5-to-0.6.md](docs/UPGRADE-0.5-to-0.6.md) — checklist, Helm move, and optional 0.6.x features.

---

## Configuration: config file, env, CLI

pgwd loads settings from (in order): **config file** → **environment variables** → **CLI flags**. Each layer overrides the previous.

| Source | Path / prefix |
|--------|---------------|
| Config file | `/etc/pgwd/pgwd.conf` (or `-config` / `PGWD_CONFIG`) |
| Environment | `PGWD_*` |
| CLI | `-flag` |

**Config file** (YAML) — keys match `-flag` and `PGWD_*` env vars. See `contrib/pgwd.conf.example` (or **`pgwd --print-sample-config > /etc/pgwd/pgwd.conf`** when the example file is not on disk). Use **`databases:`** for one or more Postgres (required even for a single target). Legacy top-level **`db:`** was removed in v1.0. For kube.postgres, use a single `databases:` entry with a URL that kube port-forward rewrites (multi-DB + kube is not supported).

```bash
# Use default path /etc/pgwd/pgwd.conf
pgwd

# Or specify path
pgwd -config /etc/pgwd/pgwd.conf
PGWD_CONFIG=/path/to/pgwd.conf pgwd
```

**Precedence:** **CLI > config file > defaults** when a config file is loaded (`PGWD_*` env vars are **not** applied in that case). When **no** config file is loaded: **CLI > env > defaults**. Use env for secrets and one-off overrides when running without a file; use a config file for stable base settings.

**`-db-url` override (one-shot):** When the config file has `databases:` (multi-DB), passing `-db-url` and `-interval 0` runs against that single URL only, ignoring the databases from config for that run. Useful for quick ad-hoc checks without editing the config.

### Multi-database limitations

- **`-kube-postgres` / `kube.postgres` and `databases:` are mutually exclusive.** Validation rejects that combination. Multi-DB mode expects **direct** Postgres URLs (e.g. in-cluster DNS, VPN, or `localhost` after manual port-forward). For port-forward from **outside** the cluster, use **N port-forwards + one `databases:` config**, or **one pgwd process per** forwarded instance — see **[docs/use-cases.md](docs/use-cases.md)** (UC-5, UC-6, UC-7).
- **Persisted history and hysteresis** (SQLite) are keyed by **`(client, cluster, database)`**. The **host from the URL is not part of that key**. If several targets share the same database name and the same **derived** client (`base client` + `-` + name from the URL path), they **collide** in the store and resolution/hysteresis will be wrong. Give each `databases:` entry a **unique `client`** when monitoring the same logical DB name on different hosts.
- **Different credentials per database:** put the full DSN (user + password) in each `databases[].url`. pgwd does not read multiple K8s Secrets in one process — inject URLs at deploy (Helm/Kustomize) or use [kubernetes-passwords.md](docs/kubernetes-passwords.md) for single-DB kube patterns.

**Operator guide (all scenarios):** **[docs/use-cases.md](docs/use-cases.md)**.

**Ready-to-use profiles:** [`contrib/profiles/`](contrib/profiles/) — `minimal-slack`, `daemon-loki`, `kube-prod`, `multi-db`.

### Anonymous usage

Daemon mode (`interval > 0`) can optionally phone home to improve pgwd (same privacy model as [gghstats](https://github.com/hrodrig/gghstats)). **Both are off by default for telemetry; update check is on unless disabled.**

| Setting | Default | Outbound call | What is sent |
|---------|---------|---------------|--------------|
| `enable_collector` / `PGWD_ENABLE_COLLECTOR` | **off** | **POST** `https://collect.gghstats.com/a1b2c3d4e5f6a7b8` (once per daemon start) | `version`, `commit`, `build_date`, one-way `hash`, boolean `features` only (e.g. multi_db, has_slack, has_loki) |
| `enable_update_check` / `PGWD_ENABLE_UPDATE_CHECK` | **on** | **GET** `https://api.github.com/repos/hrodrig/pgwd/releases/latest` | No pgwd config; public release tag only (semver compare) |

**Never sent to either destination:** DSN/URL, hostnames, database names, `client`, cluster/namespace, webhook URLs, file paths, Loki labels, or any secret.

The ingest host is [collect.gghstats.com](https://collect.gghstats.com) (shared Hermes collector; server tags reports as `project=pgwd`). Operators can audit the payload at debug log level (`log_level: debug`). Errors are debug-logged only; outbound calls never block monitoring.

**Example collector payload** (POST body; illustrative values):

```json
{
  "version": "1.2.0",
  "commit": "abc1234",
  "build_date": "2026-07-18T12:00:00Z",
  "hash": "a1b2c3d4e5f67890",
  "features": {
    "multi_db": false,
    "uses_level_mode": true,
    "long_query_enabled": false,
    "has_slack": true,
    "has_loki": true,
    "has_kube_postgres": false,
    "has_kube_loki": false,
    "has_sqlite_store": true,
    "has_sql_metrics_store": false,
    "has_http_listen": true,
    "confirm_alert_gt_1": false,
    "confirm_ok_gt_1": false,
    "dry_run": false
  }
}
```

`hash` is a short one-way fingerprint of the feature shape (dedup only; not reversible config).

### Using only environment variables

```bash
export PGWD_DB_URL="postgres://user:pass@localhost:5432/mydb"
export PGWD_DB_THRESHOLD_LEVELS="75,85,95"
export PGWD_DB_THRESHOLD_IDLE=50
export PGWD_NOTIFICATIONS_SLACK_WEBHOOK="https://hooks.slack.com/services/..."
export PGWD_INTERVAL=60

pgwd
# Runs as daemon every 60s; no need to pass any flag.
```

### Env for defaults, CLI to override

```bash
export PGWD_DB_URL="postgres://localhost:5432/mydb"
export PGWD_DB_THRESHOLD_LEVELS="70,85,90"
export PGWD_NOTIFICATIONS_SLACK_WEBHOOK="https://hooks.slack.com/..."

# Override DB and run once (e.g. for a different host)
pgwd -db-url "postgres://prod-host:5432/mydb" -interval 0

# Override threshold for a quick test
pgwd -db-threshold-levels 5,10,15 -dry-run
```

---

## Usage examples

### By threshold type

| Threshold | Use when you care about… | Example |
|-----------|---------------------------|--------|
| **levels** (3-tier) | % of `max_connections` — attention / alert / danger (default mode) | `-db-threshold-levels 75,85,95` (default) or `-db-threshold-levels 70,85,90` |
| **idle** | Pool size / connections sitting idle | `-db-threshold-idle 40` |
| **stale** | Connections open too long (leaks, never closed) | `-db-stale-age 600 -db-threshold-stale 1` |

```bash
# 3-tier levels (default 75,85,95% of max_connections) — one-shot, Slack
pgwd -db-url "postgres://user:pass@localhost:5432/mydb" \
     -notifications-slack-webhook "https://hooks.slack.com/services/..."

# Custom levels (e.g. 70,85,90%) — one-shot, Loki
pgwd -db-url "postgres://..." -db-threshold-levels 70,85,90 -notifications-loki-url "http://localhost:3100/loki/api/v1/push"

# Idle connections ≥ 40 (daemon every 60s, Slack)
pgwd -db-url "postgres://..." -db-threshold-idle 40 -interval 60 -notifications-slack-webhook "https://..."

# Stale: ≥ 1 connection open longer than 10 minutes
pgwd -db-url "postgres://..." -db-stale-age 600 -db-threshold-stale 1 -notifications-slack-webhook "https://..."
```

### Multiple thresholds in one run

You can combine several thresholds; each one that is exceeded generates an alert (same run can send multiple events).

```bash
# Alert on levels (3-tier) OR idle OR stale in a single run
pgwd -db-url "postgres://..." \
     -db-threshold-levels 75,85,95 \
     -db-threshold-idle 60 \
     -db-stale-age 600 -db-threshold-stale 1 \
     -interval 120 \
     -notifications-slack-webhook "https://..." \
     -notifications-loki-url "http://localhost:3100/loki/api/v1/push"
```

### By notifier

```bash
# Slack only (default 3-tier levels)
pgwd -db-url "postgres://..." -notifications-slack-webhook "https://hooks.slack.com/..."

# Loki only (optional labels)
pgwd -db-url "postgres://..." \
     -notifications-loki-url "http://localhost:3100/loki/api/v1/push" \
     -notifications-loki-labels "app=pgwd,env=prod,db=myapp"

# Slack and Loki (same event sent to both)
pgwd -db-url "postgres://..." \
     -notifications-slack-webhook "https://hooks.slack.com/..." \
     -notifications-loki-url "http://localhost:3100/loki/api/v1/push"
```

### Run mode and dry-run

| `interval` | Behavior |
|-------------|----------|
| **0** | One-shot — check once, then exit |
| **> 0** (e.g. 60) | Daemon — check every N seconds until Ctrl+C or SIGTERM |

**Recommendation: use daemon mode** (`interval` > 0) when you need resolution notifications, hysteresis (`confirm_ok`, `confirm_alert`), the HTTP server (`/metrics`, `/healthz`), or SQLite metrics history. In one-shot or timer mode, pgwd runs once per tick and exits — there is no state history, so resolution alerts, Prometheus scraping, and hysteresis do not apply.

```bash
# One-shot: run once, then exit (ideal for cron)
pgwd -db-url "postgres://..." -notifications-slack-webhook "https://..."
# or: PGWD_INTERVAL=0 pgwd

# Daemon: run every N seconds until Ctrl+C or SIGTERM
pgwd -db-url "postgres://..." -interval 60 -notifications-slack-webhook "https://..."

# Dry run: only print stats (total/active/idle), no notifications; no webhook/loki needed
pgwd -db-url "postgres://..." -dry-run
# With interval > 0 (default 60): runs as daemon, prints every interval — Ctrl+C to stop
# With interval 0: runs once and exits — quick connectivity test
pgwd -db-url "postgres://..." -dry-run -interval 0

# Force notification: send a test message to all configured notifiers (no threshold required)
# Use to validate delivery and format before relying on real alerts
pgwd -db-url "postgres://..." -notifications-slack-webhook "https://..." -force-notification
pgwd -db-url "postgres://..." -notifications-loki-url "http://localhost:3100/loki/api/v1/push" -force-notification
```

**Quick test** (after install): `pgwd -dry-run -interval 0` — one check, prints stats, exits. Use config file or `-db-url` + `-config`.

[↑ Back to top](#top)

---

## Typical scenarios

| Scenario | Suggestion |
|----------|------------|
| **Many Postgres instances** | Use `databases:` in one config (daemon mode) for multiple **direct** URLs. You cannot combine `databases:` with `-kube-postgres`. Use a **distinct `client` per entry** when the same DB name appears on different hosts (see [Multi-database limitations](#multi-database-limitations)). For diverse setups (different clusters, kube contexts), one config per instance with cron may be simpler. |
| **Cron check every 5 min** | One-shot (`interval` 0 or unset), one or more thresholds, Slack or Loki. Run from cron every 5 minutes. No resolution alerts, /metrics, or hysteresis. |

Frequently Asked Questions

What is pgwd?

pgwd is a Database library for the Go programming language. CLI that monitors PostgreSQL connection counts (total, active, idle, stale) and notifies via Slack and/or Loki when thresholds are exceeded. Supports Kubernetes (kubectl port-forward) and optional run context in notifications

How do I install pgwd?

Install pgwd with the Go module system using `go get hrodrig/pgwd`. Check the repository for the current installation instructions.

What category does pgwd belong to?

pgwd is listed under Database, specifically Database Tools.

← Back to Database Tools