The plan arrives as a markdown file, and it is genuinely good writing:
Move the fifty MySQL clusters to TiDB. One elastic cluster instead of fifty, horizontal scale when we need it, rolling upgrades, and we stop losing a week a month to patching. The wire protocol is MySQL-compatible, so the application layer barely changes.
The engineer who wrote it is good, and is not a database specialist, which is fine. The motivation is the most real thing in the room. Patching fifty clusters across four environments is miserable work, nobody can say with confidence which of the several hundred services depends on which cluster, and every maintenance window is a negotiation with teams who don’t reply. They read a well-argued post about someone’s TiDB move, spent fifteen minutes with a model filling in the shape of it, and wrote it up.
The meeting goes fine. The staff engineers nod. The director calls it interesting and asks for a doc. Nothing is funded, nothing is scheduled, and not one technical claim in the plan gets challenged.
That last part is the part worth explaining, because being ignored and being refuted are different outcomes and the second one is better. A refuted proposal tells you where the wall is. A proposal that draws polite agreement and no money tells you nothing, and the natural read (they’re change-averse, they didn’t get it, there’s no appetite) is almost always wrong. The plan was not rejected on its merits. It was discounted before its merits came up, and the discount is mechanical.
Nobody argued because nobody had to
A product owner once said this out loud, more plainly than any of the theory below. The question put to him was why an obviously beneficial change wasn’t getting scheduled. His answer was that I had spent no time on it, and that if I believed in it as much as I was claiming, I would have spent some. He wanted a proof of concept, something that ran. Talk is cheap, he said, and conceptually anything can sound good.
The second half of that is the part worth keeping. He wasn’t saying the idea was wrong, or even that he doubted the claim. He was reading my investment as an estimate of my own confidence, which is a better signal than anything I could have written down, and a harder one to fake. Someone who won’t spend two days on a change they say will save the company a week every month has quietly told you their real estimate of the odds, whatever the document says. Proposals get read that way whether or not the reader has the vocabulary for it.
There is a formal name for the general case. In Crawford and Sobel’s 1982 model of strategic information transmission (Econometrica 50:6), a message that is costless to send and cannot be verified carries information only to the extent that the sender’s and receiver’s interests already coincide. Economists call it cheap talk, and the result is not that listeners are hostile. A rational listener extracts close to nothing from an unverifiable claim that cost nothing to make, however well constructed. They don’t argue because there is no proposition to argue with yet, only an assertion about a future nobody has priced. The interests don’t coincide by default either: the person proposing the migration and the person on call for it in month four are usually not the same person, and everyone in the room knows it.
None of this is new, and what changed isn’t the amount of thought. The older version of this proposal was a link: read this post about how Company X moved off MySQL, with a sentence of framing on top. Fifteen minutes of investment, and everyone could see it was fifteen minutes. The current version is a twelve-page markdown file with a phased rollout, a risk table and a checklist, produced in the same fifteen minutes. The investment didn’t go up; the artifact got much better at looking like the output of a month. Which works against the author, because a reader who used to price the effort from the format can’t any more, so they stop reading the format and discount by default. The polish removed the last cheap signal the author had.
Then the harder version, which is the mechanism rather than a comment on anyone’s diligence. The plan reads cleanly because of what it leaves out, and it leaves those things out because the author cannot see them yet. Consolidating fifty clusters is easy to write while the fifty clusters are a number. It gets hard once you know that cluster 31 runs a quarterly job with no owner, that the auth service touches four of them inside one transaction, and that two environments are on a version the target won’t take. Understanding the dependencies isn’t a step the author skipped on the way to the proposal. It’s the thing that would have changed the proposal, and usually shrunk it into something smaller, duller, and fundable.
The obvious response is to make the document better: a cost-benefit table, an exec sponsor, a risk register in an appendix. It doesn’t work, because more document is still document and every line is still costless and unverifiable. Weak reasons survive that treatment perfectly well, and they share a tell: they price the layer everyone can see and treat the rest as details. When Uber published its 2016 write-up of moving from Postgres back to MySQL, the lesson was not a direction of travel but that “forward” depends on whose workload you are standing in.
The price list
Here is what the plan is actually proposing, stacked from the layer everyone estimates to the one nobody does. The pair rotates and the shape holds every time. Priced below for MySQL to Postgres, the move with the most public evidence behind it, and treat the five as the floor of an estimate rather than the estimate.
Layer one: the query surface
This layer is semantic, not syntactic. The queries that throw a syntax error are the safe ones, because the migration stops and makes you fix them. The dangerous class runs on both engines and means something different on each. InnoDB defaults to REPEATABLE READ with next-key locking and Postgres to READ COMMITTED with per-statement snapshots, so a transaction ported verbatim takes different locks, and retry logic tuned against gap-lock deadlocks meets conflicts it was never shaped for. A utf8mb4_general_ci column compares case-insensitively and a text column does not, so a UNIQUE index on email quietly changes meaning (that drift inside a single MySQL estate is its own story). Schemas predating strict mode have tolerated silent truncation and zero-dates for years, so the backfill is where you learn what the old system was accepting, on data that looks fine.
Which sets the acceptance bar: a query hasn’t ported when it runs, it’s ported when it returns the same rows. All of that fails silently, so the check is a diff, not a green test suite.
Layer two: the coupling surface
This is the layer that doesn’t move with the database. The ORM, the query builders, the raw SQL beside them, the analytics jobs on a replica: all of it grew up expecting a particular engine, and the code around the schema is where the exit gets expensive, which is its own post.
At least the application lives in a repo someone can grep. The coupling that goes unmapped never made it into version control: the crontab calling mysqldump nightly, the retention script pruning rows at 3am, the mysql -e one-liners in deploy scripts on a bastion three people know about. Reports hide best, because they run on a cadence. A month of watching the query log catches the weekly ones; the quarterly close fires long after the audit ended, and the first anyone hears is finance asking why the numbers stopped. Then the things the database feeds rather than serves, since a CDC pipeline reading the binlog gets rebuilt against logical decoding rather than repointed. This is where the two-sprint estimate quietly doubles, before a single row has moved.
Layer three: the data move
Moving the schema is easy. Moving data under continuous writes is the project: a backfill, a dual-write phase, a reconciliation process proving the two sides agree, a cutover. Dual-write is the phase that quietly becomes permanent.
Reconciliation is a subsystem with an owner, close kin to the cross-instance jobs that appear whenever two stores have to agree. Cutover is the easy-looking part, and it goes badly when the earlier phases were rushed, because it is where every deferred difference in layer one arrives at once.
Layer four: the operational reset
This layer sets the team’s intuition back to zero. A group fluent in operating MySQL is, on the day of cutover, a junior team operating Postgres. A decade of index decisions leaned on InnoDB clustering that Postgres heaps do not provide. Postgres forks a backend per connection rather than a thread, making a pooler mandatory on day one rather than an optimization for later. VACUUM, bloat and transaction-ID wraparound are an operational surface MySQL refugees have never carried, and none of it announces itself until it is a problem. None of this is unsolvable. It is a team relearning operations while carrying a pager.
Layer five: the long tail
This is the default outcome rather than the failure case. The honest steady state of a large migration is 90% done and staying there, because the last 5% is the part with no clean port, and the economics of finishing it never beat the economics of leaving the old system up for just those callers. The migration didn’t fail. It never ended.
The price, and who is left holding it
An estimate that prices all five layers looks nothing like the one in the plan. Every layer gets a line, the coupling rewrite usually the largest, the off-repo pieces inventoried rather than discovered. And the long tail gets an explicit answer: which callers move, which get deprecated, the drop-dead date for the old system, and whose name is against it.
That last item is not paperwork. It’s the only line on the page the author cannot write for free, which is what stops the proposal being cheap talk: a cost estimate is a claim about the future, while a name against the long tail costs the person making it something. Which is why the question that lands is never “have you considered the risks,” inviting another paragraph, but “who is turning the old clusters off, and when.”
Push on that and three harder ones fall out, and they decide whether you spend the next decade running both.
Is the proposer doing this work, or proposing somebody else does it? Same text, different document. A plan assigning eighteen months to a team that hasn’t agreed to it is a request, and the cheap-talk logic bites hardest here, because the author has arranged to hold none of the cost.
If it is the proposer, what happens when they leave? Eighteen months outlasts plenty of tenures, and a migration that depends on one person’s continued employment has a bus factor of one on its most expensive phase. A name is where the answer starts, not the answer. The commitment has to land on the team that carries the pager for the thing, with the work on that team’s roadmap, so the owner is a role that gets backfilled rather than a person who gets replaced by nobody.
And the question underneath both: what actually stops this becoming two systems? Nothing does, by default, which is why that is the usual result. The structural fix is the definition of done. Most migrations define done as traffic serving from the new engine, which is exactly when attention leaves, and the last five percent then never finishes because nobody’s objectives mention it. Define done as the old system decommissioned and its budget line closed, fund the decommission in the same phase as the cutover, and the tail stops being someone’s spare-time project.
Which suggests an order that tests all of it cheaply, and it runs backwards from how these are usually planned: migrate the long tail first. Take the batch job someone wrote in 2019, the integration hard-wired to the old connection string, the report with the engine-specific function, and move those before touching the easy ninety-five percent. If they move, the rest is downhill and the risk that actually kills migrations is already retired. If they can’t, the project isn’t viable at the quoted price, and that’s been established in week three rather than month fourteen.
What makes it credible
The same test applies to the premise, and here the premise breaks first. The stated reason is patching, and nobody measured the patching. How many distinct versions run across the fifty clusters, how long one cycle takes, what share of that is the database rather than the dependency negotiation: none of it is in the document, and all of it is a query or an afternoon with the change log.
The promise doesn’t survive a check in the other direction either. A TiDB cluster is not one thing to patch instead of fifty. Per PingCAP’s own architecture documentation it is a stateless SQL layer, a metadata and scheduling layer (PD), a distributed key-value store (TiKV), and optionally a columnar store for analytics. Rolling upgrades across those beat fifty negotiated windows, and it is still more component types under management, not fewer. Maybe the trade is worth it. The point is that the central claim was testable and went untested, and the room could see that from the first paragraph.
Two numbers fix that. What it costs: the five layers priced in engineer-weeks, plus the change in what the estate costs to run once it’s done. And what it buys back, in units the business already counts. Maintenance hours per month. Negotiated windows per quarter. How long one routine version bump takes end to end, today, measured rather than remembered.
The working behind them is the enumeration, and the enumeration is the deliverable. Not a risk register, which is a list of adjectives. A list with rows in it. Every service connecting to each cluster, named, with a current owner. What has to change in each and why, usually a short answer: a driver version, a connection string, a query leaning on behavior the target doesn’t implement. How many of the fifty clusters are genuinely distinct versus copies of one pattern, because that moves the project size by an order of magnitude. The versions actually running in every environment, which is the patching claim’s own evidence and nobody has it written down. And the trade-offs stated out loud, because a proposal listing no costs is reporting that none were looked for.
That list is expensive to fake, which is the point. No tool produces it in fifteen minutes, because most of it isn’t written down anywhere to retrieve; it comes from connection metadata, change logs and a dozen conversations. It also tends to answer the question it was built for. A third of the way into an honest inventory, most people find either that the project is three times what they said, or that the real win is a smaller change they can ship next quarter.
The strongest version of the same signal is the proof of concept the product owner asked for, with one condition on what it proves. A POC that stands up the target engine and shows the queries running proves layer one, and layer one was never in doubt: the compatible wire protocol is the claim everyone already believes. The slice has to run all five layers instead, on a schema with real writes and a small blast radius, ending with the old thing turned off. That is price discovery, and it comes before the funding request, being small enough to need permission rather than budget.
What actually gets funded
None of which, on its own, gets it funded, because credible and funded are different bars. The second is not short-termism, and reading it that way is the mistake that keeps these proposals coming back. A company will spend a year and real money without blinking when the year ends in more profit; it does that constantly. What it will not spend a year on is making engineering pleasanter. Easier operations for the people who run the operations is not a return. It appears nowhere in the P&L, and it converts to money only through headcount nobody intends to cut. A smaller infrastructure bill lands the same way, since eighteen months of engineer-time costs more than most database estates cost to run, which makes the ask millions to save thousands.
What does get funded is a payoff denominated in revenue, and there are two doors into that room. The first is a ceiling costing business you would otherwise take: we cannot onboard the customer who needs a region we can’t shard into, we cannot carry the volume the next contract tier brings, we said no in Q3 because the write path tops out. The second is a customer-visible change big enough to move retention or upsell, which for a database usually means latency the customer feels: the p99 that makes the product seem broken on precisely the accounts with the most data in them, the report that times out so nobody upgrades to the analytics tier, the checkout shedding orders every peak. Both dwarf five layers, and both arrive already denominated in the units the decision gets made in. Savings are not a third door.
Which settles the plan from the top of this post. Moving fifty MySQL clusters to TiDB becomes viable the day a customer the size of Walmart is onboarding and the current estate demonstrably cannot carry them. That version writes itself, survives every question here, and gets funded in a week. Without that customer, the same document asks for eighteen months to make patching less annoying, which is a real problem and the wrong ask.
The corollary is uncomfortable. Most databases called outgrown are nowhere near a ceiling and have years of headroom in unglamorous query and schema work, which is its own series. If nothing is being turned away and no customer feels it, the honest finding is that this shouldn’t be funded this year, and the engineer who works that out and says so has produced something more useful than the plan.
There is a shape problem underneath the arithmetic too. The plan is a waterfall: eighteen months in which nothing gets better, ending in a cutover that concentrates the entire risk at the last moment, and whose failure mode is not a delay but a permanent second system. Set against that, the reward is a handful of engineers having a better time patching. A large risk taken all at once for a small and diffuse gain is a trade nobody signs, however the document reads.
Which is where the smaller proposal hiding inside the big one turns up, and it has the opposite shape. The pain was patching fifty clusters by hand. The fix for that is a pipeline that patches them: a build job rolling a version through environments, a canary, health checks gating promotion, and the dependency list deciding the order. It delivers in week three and keeps delivering, each cluster onboarded to it a standalone improvement that can be stopped, reversed, or paused without leaving anything half-migrated. No dual-write, no operational reset, no second database to run forever. It clears the bar easily too, because the bar scales with the ask. Productivity is a perfectly good argument for two sprints of automation and a hopeless one for eighteen months of migration. And the enumeration that failed to justify TiDB is exactly the input the automation needs, which is why doing it first is never the wasted work it looks like.
Which is worth naming, because the three fixes above are not three flavors of trying harder. Cheap talk is defined by three properties: the message costs nothing to send, it binds nobody, and no one can check it. Each fix breaks one of them. The enumeration makes the claim checkable. The proof of concept makes it costly, which is what the product owner was asking for and could not have put in those words. The name and the date make it binding. Any one of the three moves the proposal out of the category the room was discounting before it reached the merits. And there is a sting in the checkable one, because verifiable disclosure unravels: once a few people start bringing the inventory, turning up without one stops being neutral and starts reading as having looked and not liked the answer.
