MoiToi.TECHTiDB EngineeringGuide
Custom TiDB metrics with Prometheus, VictoriaMetrics and Grafana
Updated · Andres Kepler
In short
TiDB's built-in metrics describe the components — TiDB servers, TiKV, PD. Incidents are usually about the workload: one query digest, one table's statistics, backup or replication lag, a business flow slowing down. Custom metrics turn those into time series with alerts, collected by Prometheus, kept in VictoriaMetrics or Mimir, and shown in Grafana around the questions the team actually asks.
01
What the built-in metrics cover
Every component exposes Prometheus metrics on its status endpoint — the TiDB server on port 10080, TiKV on 20180, PD on 2379 — and TiUP and TiDB Operator deploy Prometheus and Grafana with dashboards for each. They are the right baseline: QPS, latency percentiles, Region health, Raft, scheduling, resource use.
They are cluster-centric. They say a TiKV node is busy; they do not say which query, which table or which customer flow made it busy.
02
The metrics teams usually miss
- — Latency and execution count per statement digest, so a regression in one query is visible before the cluster average moves.
- — Statistics health per important table, because a drop predicts the plan change that comes next.
- — Log backup checkpoint lag as the live recovery point, and changefeed lag for TiCDC replication.
- — Hot tables and Regions over time, not only during an incident.
- — Workload signals the business cares about — orders written per minute, jobs completed — next to the database signals that explain them.
03
How to build them
Three sources cover most needs. Existing component metrics, combined with recording rules in Prometheus or vmalert. SQL-based exporters that query TiDB's system tables — statements summary, cluster information, statistics — on an interval and publish the result as metrics. And metrics from the application itself, labelled so they can be lined up with the database.
Prometheus collects; VictoriaMetrics or Mimir keep the data for months at a manageable cost and let many clusters share one store; Grafana shows it.
-- The 20 statement digests that cost the most time, cluster-wide:
-- a good source for a per-digest latency metric
SELECT SCHEMA_NAME, DIGEST, EXEC_COUNT, AVG_LATENCY, SUM_LATENCY
FROM INFORMATION_SCHEMA.CLUSTER_STATEMENTS_SUMMARY
ORDER BY SUM_LATENCY DESC
LIMIT 20;04
Keep cardinality bounded
Every distinct label combination is a new time series. Labelling by raw SQL text, user id or request id makes the series count grow without limit, and no storage engine makes that cheap. Use statement digests and a top-N, keep label values from a known set, and decide retention per metric rather than once for everything.
05
Resource groups and the million-series wall
TiDB's own metrics can cause the problem too. Enabling resource control adds a resource-group dimension to many metrics, and on a large cluster the series count multiplies with it. In production we have seen a single TiDB monitoring Prometheus pass a million active series after resource groups were enabled — and at that scale it stops coping: scrapes and rule evaluation fall behind, memory runs out, and the dashboards go blank exactly when they are needed.
Enabling resource groups should therefore be treated as a monitoring change as well as a database change: measure the series count before and after, drop or relabel the series nobody uses, aggregate per-group detail with recording rules, and move long-term storage to VictoriaMetrics or Mimir, which are built for this volume — before the rollout, not after the outage.
# Which metric names hold the most series right now (PromQL)
topk(20, count by (__name__) ({__name__=~".+"}))
# Total active series in the TSDB
prometheus_tsdb_head_series06
Alerts with a stated response
An alert that wakes someone should say what is wrong and what to do first. Each custom metric earns an alert only if there is a response to write next to it; the rest belong on a dashboard. That is the difference between monitoring that exists and monitoring that leads to decisions.
Next step