Skip to content
MOITOITECH
EET--:--:--
SUN--:--

MoiToi.TECHSpecialist practice

Observability Engineering

Prometheus. VictoriaMetrics. Mimir. Grafana. Alerts.

Monitoring that still works at scale — and leads to a decision when something breaks.


Problems

The problems this is for.

Teams rarely ask for observability. They arrive with one of these.
  • Prometheus that falls over

    Memory climbs, scrapes and rule evaluation fall behind, restarts take minutes to replay — usually because the active series count has grown past what one Prometheus can hold.

  • Cardinality nobody planned

    A new label, a feature flag or a database feature multiplies the series count overnight. Enabling TiDB resource groups is a common example: we have seen it push one Prometheus past a million series.

  • Retention you cannot afford

    Keeping months of metrics, or many clusters in one place, costs more in Prometheus than the data is worth to you.

  • Dashboards that do not answer the question

    Hundreds of panels, and still no quick way to tell what is broken, where, and what to do next.

  • Alerts that get ignored

    Too many pages, too few with a clear response, and the important one lost in the noise.

  • Kubernetes and cloud blind spots

    Workloads, nodes, databases and the application each monitored separately, with nothing that lines them up during an incident.


Packages

Start with a review, fix with a sprint.

The Review is the low-risk way in. The sprints are for a specific problem to fix and measure.

START HERE

Observability Review

2–3 focused daysFixed price, agreed before work starts

For a team that wants an independent read on its monitoring before deciding what to change.

You get

  • — series count and cardinality analysis — which metrics and labels carry the cost
  • — storage, retention and cost assessment
  • — dashboard and alert review against real incident questions
  • — written findings ranked by risk, with fix-now and fix-later recommendations

SCALE

Metrics Scale & Migration Sprint

2 weeks minimumFixed price per sprint, scoped up front

For a Prometheus that has hit its limit, or long-term metrics that need a new home.

You get

  • — cardinality reduction: dropped, relabelled and aggregated series
  • — VictoriaMetrics or Mimir architecture sized to the workload
  • — migration with both systems running in parallel and the results compared
  • — Grafana datasources switched over with nothing lost

DECISIONS

Dashboards & Alerting Sprint

2 weeks minimumFixed price per sprint, scoped up front

For monitoring that exists but does not lead to decisions.

You get

  • — dashboards built around questions, not panels
  • — alerts with a stated response, and the noisy ones removed
  • — custom metrics for the workload and the business signals that explain it
  • — capacity signals and troubleshooting workflows

The path

From telemetry to a decision.

The work covers the whole path: what is collected, where it is stored at what retention and cost, which dashboards answer which questions, and which alerts lead to which action.
  1. Exporters & app metrics
  2. Prometheus / vmagent
  3. VictoriaMetrics / Mimir
  4. Grafana
  5. Alerts
  6. Decisions

Running TiDB? The database side — custom TiDB metrics, resource groups and statement-level latency — is covered by TiDB Engineering and the guide to custom TiDB metrics.


Background

Who does the work.

Andres Kepler

Andres Kepler

Product engineer and infrastructure specialist

The engineer on the engagement, from Review to handover.

Prometheus, VictoriaMetrics, Mimir and Grafana are named as the technology the work uses. MoiToi.TECH is independent and is not a partner, reseller or endorsed provider of any monitoring vendor.


Questions

Monitoring questions, answered.

Why does Prometheus run out of memory?
Almost always because of the number of active time series. Every distinct label combination is a series held in memory, so a high-cardinality label — request ids, user ids, raw paths, or a database feature that adds a dimension — multiplies memory use. Past roughly a million active series a single Prometheus usually struggles: scrapes and rule evaluation fall behind and restarts take a long time. The fix is to find which metrics and labels carry the series, drop or aggregate them, and move long-term storage to a system built for the volume.
How do I find high-cardinality metrics in Prometheus?
Start with prometheus_tsdb_head_series for the total, then topk(20, count by (__name__) ({__name__=~".+"})) for the metric names holding the most series, and the TSDB status page for the labels with the most values. Then decide, per metric, whether anyone uses it — most large series counts come from a handful of metrics nobody looks at.
Prometheus vs VictoriaMetrics vs Mimir — which should we use?
Keep Prometheus or vmagent for collection. For long retention, many clusters or high series counts, store the metrics in VictoriaMetrics or Grafana Mimir. VictoriaMetrics is simpler to run and very resource-efficient; Mimir fits teams already invested in the Grafana stack and object storage. The choice depends on scale, retention, cost and what the team already operates — the Observability Review makes it with your numbers.
How do you migrate from Prometheus to VictoriaMetrics without losing data?
Run both in parallel. Prometheus keeps scraping and remote-writes to VictoriaMetrics; dashboards and alerts are compared on both; historical data is backfilled from Prometheus snapshots with vmctl where it is needed. Only when the results match are Grafana datasources and alerting switched, and Prometheus retention reduced.
Does enabling TiDB resource groups affect monitoring?
Yes. Resource control adds a resource-group dimension to many TiDB metrics, and on a large cluster the series count multiplies. We have seen a TiDB monitoring Prometheus pass a million active series after resource groups were enabled. Treat it as a monitoring change: measure series before and after, drop what nobody uses, aggregate with recording rules, and move storage to VictoriaMetrics or Mimir first. Guide: moitoi.tech/tidb/custom-metrics.
Our dashboards don't help during incidents. What changes?
Dashboards are rebuilt around the questions asked during an incident — what is broken, where, since when, and is it getting worse — and every alert that pages someone gets a stated response. Alerts without a response move to dashboards. Workload and business signals sit next to the infrastructure signals that explain them.
Who does the observability work?
Andres Kepler of MoiToi.TECH, an independent engineering practice in Estonia working with teams across Europe, with years of production work on Prometheus, VictoriaMetrics, Mimir and Grafana around Kubernetes, cloud and distributed databases. MoiToi is not a partner or reseller of any monitoring vendor.

Next step

Monitoring that stopped keeping up?

Thirty minutes to describe it. You will get an honest view of whether a Review, a sprint or nothing at all is the right next step.