MoiToi.TECHTiDB EngineeringGuide
TiDB backup, restore and point-in-time recovery
Updated · Andres Kepler
In short
TiDB point-in-time recovery combines two things: BR snapshot backups of the whole cluster and a continuous log backup of every change since. A PITR restore loads the latest snapshot before the target time and replays the log up to it. A backup only counts once a full restore has been run and timed against the recovery time the business assumes.
01
The backup tools
- — BR snapshot backup — a consistent, physical backup of the whole cluster (or chosen databases and tables) at one timestamp, written in parallel by the TiKV nodes to object storage such as S3, GCS or Azure Blob.
- — BR log backup — a continuous stream of changes to the same kind of storage, started once and left running. Together with snapshots it is what makes point-in-time recovery possible.
- — Dumpling — a logical export to SQL or CSV. Useful for migrations, audits and small datasets; not a fast way to recover a large cluster.
- — On Kubernetes, TiDB Operator wraps these as Backup, BackupSchedule and Restore resources, and on some clouds can also take volume-snapshot backups.
02
How point-in-time recovery works
To restore to a moment — just before a bad deploy or a mistaken DELETE — BR restores the most recent snapshot taken before that moment and then replays the log backup up to the chosen timestamp. The restore normally goes to a fresh, empty cluster, which is then verified and switched to.
The recovery point is therefore bounded by how current the log backup is, and the recovery time by how fast the target cluster can ingest the snapshot and replay the log.
# Start continuous log backup once; it keeps running
br log start --task-name=pitr --pd "<pd>:2379" --storage "s3://backup/log"
# Restore to a point in time from snapshot + log
br restore point --pd "<pd>:2379" \
--full-backup-storage "s3://backup/snapshot-2026-10-01" \
--storage "s3://backup/log" \
--restored-ts "2026-10-03 14:05:00+0300"03
What to monitor
- — Log backup checkpoint lag — how far behind the present the log backup is. This is your real recovery point, and it should alert.
- — Log backup task status — a paused or failed task silently stops the PITR window from growing.
- — Snapshot backup success and duration — on a schedule, with retention matched to how far back you need to go.
- — Storage retention and cost — snapshots and logs older than the recovery window cost money and protect nothing.
04
Proving the backup
The usual failure is not a missing backup. It is a backup that has never been restored, so nobody knows whether it works or how long it takes. A restore drill answers both: restore to a scratch cluster on a schedule, check the data, and record how long it took against the recovery time objective (RTO) and the checkpoint lag against the recovery point objective (RPO).
Next step