Skip to content
MOITOITECH
EET--:--:--
SUN--:--

MoiToi.TECHSpecialist practice

TiDB Engineering

Migration. Performance. Reliability. Observability.

Hands-on TiDB engineering for teams running serious production workloads.


Problems

The problems this is for.

Teams rarely arrive asking for TiDB expertise. They arrive with one of these.
  • A migration you cannot afford to get wrong

    Moving a live workload onto TiDB, between TiDB clusters, or off one — with the cutover, not the data copy, being the part that keeps people awake.

  • Latency that nobody can explain

    Slow queries, hotspots, uneven regions or a cluster that degrades under load, and no agreed account of why.

  • Dashboards that do not answer the question

    Metrics exist, but when something goes wrong they do not tell you what is happening or what to do next.

  • Capacity decided by guesswork

    Is it time to scale, rebalance or change the workload — and how would you know before it hurts?

  • Backups nobody has restored

    Backup, restore and point-in-time recovery that exist on paper but have not been proven against the recovery time the business assumes.

  • Incidents that repeat

    The same class of incident returning, operational toil around the cluster, and infrastructure that drifted away from its code.


Packages

Start small, with a fixed scope.

The Health Check is the low-risk way in. The sprints are for when there is a specific problem to solve and measure.

START HERE

TiDB Health Check

2–3 focused daysFixed price, agreed before work starts

For a team that wants an independent read on its cluster before deciding what to spend on.

Answers

  • — Is the cluster healthy?
  • — Where are the obvious risks and bottlenecks?
  • — Is the observability good enough to run it?
  • — Are the backup and recovery assumptions credible?
  • — What should be fixed now, and what can wait?

You get

  • — written findings, ranked by risk
  • — fix-now / fix-later recommendations
  • — a read-out call with your engineers

PERFORMANCE

TiDB Performance Sprint

2 weeks minimumFixed price per sprint, scoped up front

For a real performance or reliability problem that needs evidence, a change and a measurement.

Answers

  • — Which workload and which queries carry the cost?
  • — Is the bottleneck the cluster, the schema or the application?
  • — What would scaling actually buy?

You get

  • — workload and query analysis
  • — bottleneck and resource analysis
  • — evidence-based configuration, schema or application changes
  • — telemetry, dashboard and alert gaps closed
  • — a before/after measurement wherever one can be taken

MIGRATION

TiDB Migration Sprint / Project

Discovery first, then a scoped sprint or projectQuoted after discovery — no two migrations are the same

For a team planning or executing a production database move.

Answers

  • — Which data movement and replication path fits this source and target?
  • — What will break on compatibility, and what will be slower?
  • — How does traffic move, how is it verified, and how is it rolled back?

You get

  • — source and target assessment
  • — migration architecture and replication choice
  • — ProxySQL cutover and rollback design
  • — data and performance validation
  • — observability before, during and after
  • — runbooks and handover

OBSERVABILITY

TiDB Observability Engineering

Scoped sprintFixed price per sprint, scoped up front

For a team whose monitoring exists but does not lead to decisions.

Answers

  • — Which signals actually describe the health of this cluster?
  • — Where should metrics live, and at what retention and cost?
  • — Which alerts should wake someone, and what do they do then?

You get

  • — telemetry and metrics-pipeline architecture
  • — dashboards built around questions, not panels
  • — alerts with a stated response
  • — capacity signals
  • — troubleshooting workflows

Migration

The cutover is the product.

Copying the data is a solved problem with well-documented tools. The dangerous part of a live migration is moving production traffic: deciding the moment, switching it, proving it worked and being able to go back. That part is what MoiToi engineers and automates.
  1. Before

    Application to ProxySQL to Source cluster

  2. During

    Data movement, replication and validation — the tool that fits the path

  3. After

    Application to ProxySQL to Target cluster

What the traffic layer controls

Pre-cutover validation
Replication lag, data checks and readiness gates that must pass before anything moves.
Hostgroup transition
Backends moved between ProxySQL hostgroups as an explicit, recorded state change.
Controlled switching
Traffic moved deliberately, with connection and routing behaviour understood in advance.
Post-cutover verification
Errors, latency and data checked against the baseline taken before the switch.
Rollback path
A tested way back, defined before the cutover rather than improvised during it.
Observability throughout
The transition is visible on dashboards while it happens, not reconstructed afterwards.

Migration paths

MySQL → TiDB

Initial workflow

The first and most proven path: moving MySQL workloads that have outgrown a single primary.

TiDB → TiDB

Next automation path

Cluster replacement and re-platforming, moves between environments and consolidation — operational migrations between TiDB clusters.

The cutover tooling is MoiToi’s own, written new against public MySQL, TiDB and ProxySQL interfaces. It grows engagement by engagement: each migration funds the automation it needed, and the generic parts carry forward to the next one.


Observability

From telemetry to a decision.

The work covers the whole path: which signals the database exposes, how they are collected and stored at a sensible retention and cost, which dashboards answer which questions, and which alerts lead to which action.
  1. TiDB telemetry
  2. Prometheus
  3. VictoriaMetrics / Mimir
  4. Grafana
  5. Alerts
  6. Operational decisions

How it works

Engineering judgement, sold as outcomes.

Outcomes, not hours
Fixed-scope packages and sprints with a definition of done — not open-ended staff augmentation.
Every sprint has a reason to exist
A sprint ends with the problem measured and handed over. The aim is your team’s independence, not a standing dependency.
Evidence before change
Recommendations come with the data behind them, and a measurement afterwards wherever one can be taken.
A clear ladder
Health Check → a Performance, Migration or Observability sprint → larger implementation only where it is earned.

Background

Who does the work.

Four years of hands-on production work with TiDB and MySQL — and with the Kubernetes, cloud, infrastructure-as-code and observability that surround a database in production. The engagement is with the person who has operated the system, not a reseller of someone else’s.
Andres Kepler

Andres Kepler

Product engineer and infrastructure specialist

The engineer on the engagement, from Health Check to handover.

Databases
TiDB · MySQL
Traffic & cutover
ProxySQL
Platform
Kubernetes · AWS · Terraform / IaC · CI/CD
Observability
Prometheus · VictoriaMetrics · Mimir · Grafana · Alerting
Practice
SRE / DBRE · Performance analysis · Capacity planning · Backup / restore / PITR · Incident analysis

MoiToi.TECH is independent. No current or former employer is a client of this practice or endorses it, and no employer’s code, dashboards, configurations, runbooks or documents are used in it. Every tool and template is built new, from first principles and public documentation.


Next step

Have a difficult TiDB problem?

Thirty minutes to describe it. You will get an honest view of whether a Health Check, a sprint or nothing at all is the right next step.