Our TiDB DBAs manage PD, TiKV, and TiFlash around the clock: continuous monitoring, P1 response under 15 minutes, rehearsed 8.5 LTS upgrades, and verified restore drills. Credgenics cut query latency from 200 ms to 50 ms with our dedicated team.
Credgenics query latency after TiKV tuning and a version upgrade.
Yulu storage on a 3 TiKV, 3 PD, 2 TiDB cluster.
CometChat moved to TiDB with zero revenue interruption.
On AWS, GCP, Azure or your own racks, managed with TiUP or TiDB Operator, and nobody on staff owns it full time.
The cutover went fine. Now hot regions, GC and PD scheduling are problems your MySQL DBAs have not met before.
Both leave maintenance in December 2026 and the upgrade path needs rehearsal, not a weekend gamble.
TiDB skills are scarce, and one engineer who knows the cluster is a single point of failure. A team covers the nights and the leave.
Operational ownership of your cluster, not a ticket queue. Every item below runs on a schedule agreed in the first 30 days.
Prometheus and Grafana alerting across PD, TiKV, TiDB and TiFlash, tuned to your workload so pages mean something. A TiDB DBA acknowledges P1 alerts inside 15 minutes, day or night.
We watch leader and region balance, split hot regions, set placement rules, and fix write hotspots at the schema level with AUTO_RANDOM or SHARD_ROW_ID_BITS when monotonic keys pile writes onto one store. Read: TiDB Scheduling.
Slow query review, execution plan binding, raft and gRPC thread tuning, region size, flow control and TiFlash replica sizing for analytical queries. This is the work that took Credgenics from 200 ms to 50 ms. Read: Optimistic and Pessimistic Locking in TiDB.
Rolling TiUP or Operator upgrades, rehearsed on a copy of production, with version pinning, a DDL freeze during the window and a rollback point. Read: What's New in TiDB 8.5.
BR snapshots plus log backup for point-in-time recovery, Dumpling exports where you need logical copies, and scheduled restores to a separate cluster so recovery time is measured, not assumed. Read: Dumpling; Understanding Flashback in TiDB.
Store usage forecasts, GC safepoint checks, TTL policies, and right-sizing of nodes and provisioned IOPS. At Credgenics, IOPS analysis cut provisioned capacity by 33%. Read: TiDB Garbage Collection.
“We appreciated the supportive and cooperative approach of the Mydbops team, especially during the TiDB upgrade process. The 75% reduction in query latency has had a significant, positive impact on our business.”
Severity-based response under a formal SLA. Real TiDB DBAs, not a ticket queue. Definitions are agreed with your team during onboarding.
| Severity | Example | First response (SLA) | Channel |
|---|---|---|---|
| P1 · Critical | Cluster unavailable, Region unavailable errors, data at risk | P1 - <15 mins, 24×7 | Phone page + war room (Google Meet or Zoom) |
| P2 · High | Latency spike, TiKV store down with replicas healthy, PD leader flapping | P2 - 30 mins | Slack or Google Chat + ticket |
| P3 · Medium | Single slow query class, TiFlash replica lag, disk trending full | P3 - 60 mins | Ticket |
| P4 · Low | Config change, user grants, planned maintenance | P4 - 90 mins | Ticket |
Minute 0: the alert fires and pages the on-call TiDB DBA. By minute 15: engaged, war room open, your team informed. Then we contain - shed load, move leaders, isolate the store - restore service, and post updates at an agreed interval until resolved. Within five working days: a written root-cause report with the fix and the prevention step.
Talk to a TiDB DBA →Monotonic keys send every insert to the same region. Fixed with key design and pre-split regions.
Block cache, coprocessor and write buffer limits reviewed against real load before the kernel kills the store.
PD quorum, leader placement and network checked; scheduling restored without guesswork.
Long transactions and TiCDC changefeeds holding GC found and cleared.
Scheduler limits and store weights adjusted so new TiKV nodes take load.
Known upgrade pitfalls, such as TiFlash replica changes or DDL mid-upgrade, are frozen out of the window.
TiDB 6.5 LTS follows on 29 December 2026. After that, fixes land only in newer LTS lines. The move to 8.5 LTS is well worth it, but upgrade bugs are real: recent PingCAP issues include clusters going down when TiFlash replicas change mid-upgrade and index inconsistency when DDL runs during an upgrade. We check OS compatibility, rehearse on a copy, and run the rolling upgrade with a rollback point.
Published Mydbops TiDB engagements - named clients, real numbers.
| Client | Problem | What we did | Result |
|---|---|---|---|
CredgenicsFintech · 98M+ loan accounts
|
200 ms dashboard latency, I/O spikes, TiKV congestion, storage at 70% | Audit, version upgrade, TiKV region and raft/gRPC thread tuning, fourth TiKV node, PD upgrade, IOPS right-sizing | 4× faster queries (200 to 50 ms), 30% CPU efficiency, 33% lower IOPS cost |
YuluMobility · 45,000+ e-bikes
|
18 TB of IoT data, write throttling, locking DDL, single master | 3 TiKV, 3 PD, 2 TiDB cluster, DM live replication, online DDL | 72% less storage (18 to 3.4 TB), no downtime during DDL |
CometChatSaaS communication
|
1.7M tables stalling MySQL during spikes | Audit, parallel TiDB build, controlled cutover, failover and tuning afterwards | 30% lower TCO, 50% storage saved, zero revenue interruption |
“Their technical depth, responsiveness, and round-the-clock support consistently stand out.”
“It's really impressive to see how Mydbops helped reduce our 18TB dataset down to just 3.4TB.”
“Working around the clock, we need help 24 hours a day, seven days a week. Mydbops efficiently provided this at all times.”
| Week | What happens | You get |
|---|---|---|
| 1 | Read-only access, topology and version review, alert audit, backup check | Baseline health report and risk list |
| 2 | Alert rules tuned, runbooks written for your cluster, escalation contacts agreed | On-call live with the P1 to P4 matrix (P1 - <15 mins, P2 - 30 mins, P3 - 60 mins, P4 - 90 mins) |
| 3 | First restore drill, slow query and hotspot review | Measured recovery time, tuning backlog |
| 4 | Upgrade and capacity plan agreed | Roadmap for the next two quarters |
| In-house DBA | PingCAP support subscription | TiDB Cloud | Mydbops remote DBA | |
|---|---|---|---|---|
| Who runs the cluster | Your hire | You; PingCAP answers tickets | PingCAP, on its platform | Mydbops, in your account |
| 24×7 coverage | Needs 3 to 4 people | By support plan | Included | Included, P1 - <15 mins |
| Keeps your infrastructure | Yes | Yes | No | Yes |
| Upgrades and backups done for you | If staffed | No | Yes | Yes, rehearsed |
| Best when | TiDB is your core product | You have strong in-house ops | You can move off self-hosting | You self-host and need ownership now |
Monthly retainer, sized by cluster count and coverage, with customisable contracts. We map your severity profile and on-call gaps to the right coverage during scoping.
Version, node count and your worst week this quarter is enough to start. A TiDB DBA replies within one business day.
Certified TiDB DBAs · under-15-minute P1 response · ISO 27001 & ISO 9001 · 800+ clients