TiDB Remote DBA Services for Self-Hosted Clusters That Cannot Stall

Our TiDB DBAs manage PD, TiKV, and TiFlash around the clock: continuous monitoring, P1 response under 15 minutes, rehearsed 8.5 LTS upgrades, and verified restore drills. Credgenics cut query latency from 200 ms to 50 ms with our dedicated team.

REVIEWED ON Clutch
★★★★★
4.9 · Verified reviews
24×7Cluster operations
P1 - <15 minsIncident response
10+Years of operations
6000+Servers managed
200 ms → 50 ms

Credgenics query latency after TiKV tuning and a version upgrade.

18 TB → 3.4 TB

Yulu storage on a 3 TiKV, 3 PD, 2 TiDB cluster.

1.7M tables

CometChat moved to TiDB with zero revenue interruption.

Trusted & certified PingCAP Service Partner ISO 27001 ISO 9001 AWS Advanced Tier Services Partner PCI DSS 800+ clients 6000+ servers under management
Who this is for

Built for teams already running TiDB in production

You self-host TiDB

On AWS, GCP, Azure or your own racks, managed with TiUP or TiDB Operator, and nobody on staff owns it full time.

You just migrated from MySQL

The cutover went fine. Now hot regions, GC and PD scheduling are problems your MySQL DBAs have not met before.

You are on 7.5 or 6.5 LTS

Both leave maintenance in December 2026 and the upgrade path needs rehearsal, not a weekend gamble.

Your on-call is one person

TiDB skills are scarce, and one engineer who knows the cluster is a single point of failure. A team covers the nights and the leave.

What we run every day

What your TiDB remote DBA team runs every day

Operational ownership of your cluster, not a ticket queue. Every item below runs on a schedule agreed in the first 30 days.

24×7 monitoring and incident response

Prometheus and Grafana alerting across PD, TiKV, TiDB and TiFlash, tuned to your workload so pages mean something. A TiDB DBA acknowledges P1 alerts inside 15 minutes, day or night.

Prometheus · Grafana · P1 - <15 mins

PD scheduling, regions and hotspots

We watch leader and region balance, split hot regions, set placement rules, and fix write hotspots at the schema level with AUTO_RANDOM or SHARD_ROW_ID_BITS when monotonic keys pile writes onto one store. Read: TiDB Scheduling.

Placement rules · hot-region splits

TiKV and TiFlash performance

Slow query review, execution plan binding, raft and gRPC thread tuning, region size, flow control and TiFlash replica sizing for analytical queries. This is the work that took Credgenics from 200 ms to 50 ms. Read: Optimistic and Pessimistic Locking in TiDB.

Plan binding · raft/gRPC tuning

Upgrades and patching

Rolling TiUP or Operator upgrades, rehearsed on a copy of production, with version pinning, a DDL freeze during the window and a rollback point. Read: What's New in TiDB 8.5.

Rehearsed · pinned · rollback point

Backup, PITR and restore drills

BR snapshots plus log backup for point-in-time recovery, Dumpling exports where you need logical copies, and scheduled restores to a separate cluster so recovery time is measured, not assumed. Read: Dumpling; Understanding Flashback in TiDB.

BR + log backup · measured RTO

Capacity, GC and cost

Store usage forecasts, GC safepoint checks, TTL policies, and right-sizing of nodes and provisioned IOPS. At Credgenics, IOPS analysis cut provisioned capacity by 33%. Read: TiDB Garbage Collection.

GC safepoints · IOPS right-sizing

“We appreciated the supportive and cooperative approach of the Mydbops team, especially during the TiDB upgrade process. The 75% reduction in query latency has had a significant, positive impact on our business.”

Naveen Malhotra - Credgenics
Naveen Malhotra Database Administrator and Architect, Credgenics
Incident response

When something breaks, here is the clock

Severity-based response under a formal SLA. Real TiDB DBAs, not a ticket queue. Definitions are agreed with your team during onboarding.

Severity Example First response (SLA) Channel
P1 · Critical Cluster unavailable, Region unavailable errors, data at risk P1 - <15 mins, 24×7 Phone page + war room (Google Meet or Zoom)
P2 · High Latency spike, TiKV store down with replicas healthy, PD leader flapping P2 - 30 mins Slack or Google Chat + ticket
P3 · Medium Single slow query class, TiFlash replica lag, disk trending full P3 - 60 mins Ticket
P4 · Low Config change, user grants, planned maintenance P4 - 90 mins Ticket

The first hour of a P1

Minute 0: the alert fires and pages the on-call TiDB DBA. By minute 15: engaged, war room open, your team informed. Then we contain - shed load, move leaders, isolate the store - restore service, and post updates at an agreed interval until resolved. Within five working days: a written root-cause report with the fix and the prevention step.

Talk to a TiDB DBA →
Failure modes we handle

The TiDB problems we get called for

Write hotspots on one TiKV store

Monotonic keys send every insert to the same region. Fixed with key design and pre-split regions.

TiKV memory climbing to OOM

Block cache, coprocessor and write buffer limits reviewed against real load before the kernel kills the store.

“PD server timeout” and unavailable regions

PD quorum, leader placement and network checked; scheduling restored without guesswork.

GC safepoint stuck, disk not released

Long transactions and TiCDC changefeeds holding GC found and cleared.

Regions not rebalancing after scale-out

Scheduler limits and store weights adjusted so new TiKV nodes take load.

Upgrades that break the cluster

Known upgrade pitfalls, such as TiFlash replica changes or DDL mid-upgrade, are frozen out of the window.

Maintenance deadline

TiDB 7.5 LTS leaves maintenance on 1 December 2026

TiDB 6.5 LTS follows on 29 December 2026. After that, fixes land only in newer LTS lines. The move to 8.5 LTS is well worth it, but upgrade bugs are real: recent PingCAP issues include clusters going down when TiFlash replicas change mid-upgrade and index inconsistency when DDL runs during an upgrade. We check OS compatibility, rehearse on a copy, and run the rolling upgrade with a rollback point.

7.5 LTS · EOM 1 Dec 2026 6.5 LTS · EOM 29 Dec 2026
Case studies

TiDB clusters we already look after

Published Mydbops TiDB engagements - named clients, real numbers.

Client Problem What we did Result
CredgenicsFintech · 98M+ loan accounts
200 ms dashboard latency, I/O spikes, TiKV congestion, storage at 70% Audit, version upgrade, TiKV region and raft/gRPC thread tuning, fourth TiKV node, PD upgrade, IOPS right-sizing 4× faster queries (200 to 50 ms), 30% CPU efficiency, 33% lower IOPS cost
YuluMobility · 45,000+ e-bikes
18 TB of IoT data, write throttling, locking DDL, single master 3 TiKV, 3 PD, 2 TiDB cluster, DM live replication, online DDL 72% less storage (18 to 3.4 TB), no downtime during DDL
CometChatSaaS communication
1.7M tables stalling MySQL during spikes Audit, parallel TiDB build, controlled cutover, failover and tuning afterwards 30% lower TCO, 50% storage saved, zero revenue interruption

“Their technical depth, responsiveness, and round-the-clock support consistently stand out.”

Ravi Ranjan - CometChat
Ravi Ranjan VP Engineering, CometChat

“It's really impressive to see how Mydbops helped reduce our 18TB dataset down to just 3.4TB.”

ND
Naveen Dachuri Co-Founder and CTO, Yulu

“Working around the clock, we need help 24 hours a day, seven days a week. Mydbops efficiently provided this at all times.”

Sunil Kumar - Shiprocket
Sunil Kumar CTO, Shiprocket (MySQL client)
First 30 days

How we take over a running cluster

Week What happens You get
1 Read-only access, topology and version review, alert audit, backup check Baseline health report and risk list
2 Alert rules tuned, runbooks written for your cluster, escalation contacts agreed On-call live with the P1 to P4 matrix (P1 - <15 mins, P2 - 30 mins, P3 - 60 mins, P4 - 90 mins)
3 First restore drill, slow query and hotspot review Measured recovery time, tuning backlog
4 Upgrade and capacity plan agreed Roadmap for the next two quarters
Decision table

In-house, PingCAP, TiDB Cloud or a remote DBA team?

In-house DBA PingCAP support subscription TiDB Cloud Mydbops remote DBA
Who runs the cluster Your hire You; PingCAP answers tickets PingCAP, on its platform Mydbops, in your account
24×7 coverage Needs 3 to 4 people By support plan Included Included, P1 - <15 mins
Keeps your infrastructure Yes Yes No Yes
Upgrades and backups done for you If staffed No Yes Yes, rehearsed
Best when TiDB is your core product You have strong in-house ops You can move off self-hosting You self-host and need ownership now
Engagement model

A retainer sized to your cluster

Monthly retainer, sized by cluster count and coverage, with customisable contracts. We map your severity profile and on-call gaps to the right coverage during scoping.

Always included

24×7 monitoring
P1 to P4 response under a formal SLA (P1 - <15 mins, P2 - 30 mins, P3 - 60 mins, P4 - 90 mins)
Dedicated customer success manager
Slack or Google Chat channel
Monthly health and security reports
Unlimited operational requests within scope

Not included

Application code changes
Schema design for new features
New cluster builds and migrations (run as TiDB consulting projects)
Cloud infrastructure outside the database layer
Common questions

Questions teams ask before they sign

We watch PD, TiKV, TiDB and TiFlash health through Prometheus and Grafana alerts, respond to incidents, review slow queries and hot regions, confirm backups and GC progress, plan capacity, and work a tuning and upgrade backlog agreed with your team. You receive a monthly health report and a security report.
P1 (cluster down or data at risk): P1 - <15 mins, 24×7. P2 (serious degradation): P2 - 30 mins. P3 (partial impact): P3 - 60 mins. P4 (requests and questions): P4 - 90 mins. A P1 opens a war room on Google Meet or Zoom with a TiDB DBA engaged until service is restored.
TiDB 7.5 LTS reaches end of maintenance on 1 December 2026 and 6.5 LTS on 29 December 2026. Most teams should move to 8.5 LTS. We check operating system compatibility first (newer releases dropped CentOS 7), rehearse the upgrade on a copy of production, pin TiUP versions, run a rolling upgrade and keep a rollback point.
TiUP rolling upgrades restart one component at a time, so most upgrades keep the cluster serving traffic. We do not promise zero impact. We rehearse on a copy of production, freeze DDL and TiFlash replica changes during the window, and agree a rollback point with you before we start.
We run BR snapshot backups with log backup for point-in-time recovery to S3-compatible storage, use Dumpling for logical exports where needed, and schedule restore drills to a separate cluster. The restore time from each drill goes into your monthly report.
Yes. We operate clusters deployed with TiDB Operator on Kubernetes as well as TiUP-managed clusters on virtual machines or bare metal, on AWS, GCP, Azure or your own data centre.
PingCAP support answers tickets about the software. TiDB Cloud runs TiDB for you on its own platform. A remote DBA runs your self-hosted cluster inside your own account: monitoring, on-call, tuning, upgrades and backups. Many teams keep a PingCAP subscription and use Mydbops for day-to-day operations.
We start with read-only monitoring access. Operational access runs through your bastion or VPN with named accounts, MFA, least-privilege database users and session logging. Break-glass access for P1 incidents is agreed in advance. Our processes are ISO 27001 certified.
You keep every runbook, alert rule, dashboard and document we built. We hand over in a recorded session, remove our accounts and help you rotate credentials.
Application code changes, schema design for new features, new cluster builds and migrations (those run as TiDB consulting projects), and cloud infrastructure outside the database layer.
Let's talk

Tell us about your TiDB cluster

Version, node count and your worst week this quarter is enough to start. A TiDB DBA replies within one business day.

Certified TiDB DBAs · under-15-minute P1 response · ISO 27001 & ISO 9001 · 800+ clients