Skip to content
MOITOITECH
EET--:--:--
SUN--:--

Run & Scale · Platform, observability, recovery

Know what will fail before it does.

Find the reliability constraint that matters to the business, then fix or prove the highest-risk path.

One entry review. One focused sprint. Three connected tracks—not a technology shopping list.

Discuss your platform

Fit

For production systems with an operational question.

This is for a CTO or platform lead with a running system and a specific risk, incident pattern, or reliability decision to resolve.
  • A Kubernetes, cloud, or infrastructure change is difficult to make safely.
  • The team has telemetry but still struggles to diagnose incidents.
  • A database restore or point-in-time recovery has not been demonstrated.
  • Manual operational work keeps recurring and needs an evidence-based fix.

Tracks

Three paths through the same reliability problem.

Start where the evidence points. Platform, observability, and recovery often meet at the same operational failure.

Platform

Deployments, infrastructure, or Kubernetes operations feel fragile or hard to change.

Trace the path from change to production. Improve the highest-risk automation, infrastructure, deployment, or rollback path that evidence supports.

Observability

Dashboards exist, but they do not help someone decide what is broken or what to do next.

Connect workload signals to collection, storage, alerts, and operational decisions. Use Prometheus-compatible tooling, VictoriaMetrics or Mimir, and Grafana where they fit.

Database recovery

Backups exist, but the recovery path or recovery objective has not been demonstrated.

Review backup assumptions, restore procedures, point-in-time recovery, observability, and runbooks. A restore is tested only when the customer environment and access permit it.


Engagement

Review first. Sprint only where it earns its place.

The existing SRE / Database Reliability Review is the entry point; implementation is a separate fixed scope.

Entry review

SRE / Database Reliability Review

For production systems where database, cloud infrastructure, observability or reliability is creating operational risk. I diagnose the system and leave you with concrete priorities — then I can stay to implement them if useful.

from €1,500 · focused engagement

A scoped assessment, not an open-ended managed SRE service or staff-augmentation contract.

Follow-on

Two-week minimum sprint

A fixed-scope intervention on the highest-value agreed platform, observability, or recovery gap. The outcome, access, and definition of done are agreed before work starts.

A sprint is not ongoing managed SRE or staff augmentation. Another sprint needs a separate outcome.


Proof

Evidence from systems that run.

Mervare is Andres’s own live pilot; Zoovet is a client system. Those are different kinds of evidence, and neither is represented as a reliability case study with measured customer outcomes.
Mervare's public harbour map

Mervare’s public system links a directory, bookings, payment integration, data pipelines, and an instrument bridge. The case study labels what is a live pilot, built, or not yet live.

Inspect the Mervare case

The existing SRE / Database Reliability Review is priced from €1,500; sprint scopes are agreed separately.