- Why does Prometheus run out of memory?
- Almost always because of the number of active time series. Every distinct label combination is a series held in memory, so a high-cardinality label — request ids, user ids, raw paths, or a database feature that adds a dimension — multiplies memory use. Past roughly a million active series a single Prometheus usually struggles: scrapes and rule evaluation fall behind and restarts take a long time. The fix is to find which metrics and labels carry the series, drop or aggregate them, and move long-term storage to a system built for the volume.
- How do I find high-cardinality metrics in Prometheus?
- Start with prometheus_tsdb_head_series for the total, then topk(20, count by (__name__) ({__name__=~".+"})) for the metric names holding the most series, and the TSDB status page for the labels with the most values. Then decide, per metric, whether anyone uses it — most large series counts come from a handful of metrics nobody looks at.
- Prometheus vs VictoriaMetrics vs Mimir — which should we use?
- Keep Prometheus or vmagent for collection. For long retention, many clusters or high series counts, store the metrics in VictoriaMetrics or Grafana Mimir. VictoriaMetrics is simpler to run and very resource-efficient; Mimir fits teams already invested in the Grafana stack and object storage. The choice depends on scale, retention, cost and what the team already operates — the Observability Review makes it with your numbers.
- How do you migrate from Prometheus to VictoriaMetrics without losing data?
- Run both in parallel. Prometheus keeps scraping and remote-writes to VictoriaMetrics; dashboards and alerts are compared on both; historical data is backfilled from Prometheus snapshots with vmctl where it is needed. Only when the results match are Grafana datasources and alerting switched, and Prometheus retention reduced.
- Does enabling TiDB resource groups affect monitoring?
- Yes. Resource control adds a resource-group dimension to many TiDB metrics, and on a large cluster the series count multiplies. We have seen a TiDB monitoring Prometheus pass a million active series after resource groups were enabled. Treat it as a monitoring change: measure series before and after, drop what nobody uses, aggregate with recording rules, and move storage to VictoriaMetrics or Mimir first. Guide: moitoi.tech/tidb/custom-metrics.
- Our dashboards don't help during incidents. What changes?
- Dashboards are rebuilt around the questions asked during an incident — what is broken, where, since when, and is it getting worse — and every alert that pages someone gets a stated response. Alerts without a response move to dashboards. Workload and business signals sit next to the infrastructure signals that explain them.
- Who does the observability work?
- Andres Kepler of MoiToi.TECH, an independent engineering practice in Estonia working with teams across Europe, with years of production work on Prometheus, VictoriaMetrics, Mimir and Grafana around Kubernetes, cloud and distributed databases. MoiToi is not a partner or reseller of any monitoring vendor.