Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Mercari rewrites its TiDB Cloud autoscaling rules by hour, weekday and campaign

Mercari's DBRE team overrides its TiDB autoscaling rules by time of day and campaign, and a cluster with a 50% CPU target now averages 44%. Extending the same controller to TiKV storage meant designing around a TiDB Cloud lock that refuses a second scale operation while one is running.

The Engineer · Build desk

How we use AISend a correction

Photograph accompanying Mercari rewrites its TiDB Cloud autoscaling rules by hour, weekday and campaign
Photo: pingcap.co.jp

What happened

  • Mercari considered scaling from the load seen a day or a week earlier and chose fixed time-window overrides so operators can predict behavior from the settings.
  • No individual node in the 50%-target cluster exceeded 80% CPU, and Mercari says its capacity goal for the stateless TiDB tier is broadly met.
  • Mercari had said in April it would pursue vertical TiKV scaling, then put horizontal scaling first after observing load during campaigns.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Anyone autoscaling TiKV on TiDB Cloud has to commit each cluster to horizontal or vertical scaling, and the right pick depends on how much data sits on how many nodes.
  • cost Campaign protection depends on someone entering the start time as a window ahead of the launch; a launch nobody scheduled runs on the default parameters.
  • capability Known traffic patterns go into readable config, so operators can predict the controller's behavior without trusting a forecast built from past load.

The controller underneath is the one Mercari described in April [1]. It reads metrics such as CPU utilization, decides how many TiDB nodes to add or remove, and makes the change through the TiDB Cloud API [2]. Its guardrails follow the Kubernetes autoscaler: a cap on nodes per step, an interval between changes, and a stabilization window that stops it alternating between scale-in and scale-out [3]. Mercari wanted daytime scale-in to stay cautious. After a certain hour at night, though, traffic fell sharply and the controller shed nodes too slowly [4]. The fix layers time-window overrides, keyed by hour and weekday, on top of the metric decision [5]. Observed load still triggers every action. The override only changes how readily the controller acts on it [5]. In the night window the cooldown is shorter and more nodes can go in one step [6]. Where load barely moves, a longer stabilization window suppresses flapping, and some windows forbid scale-in outright [8]. Weekday and hour combine because Mercari's weekday load tends to climb from around noon, while holidays show no comparable change at that hour [9]. I'd make the same choice of a schedule over a replay of past load for a service whose heaviest days sit on a campaign calendar [10]. A scaler that copies yesterday's or last week's curve does well on ordinary days. It does worst on the day that differs from history, and a campaign launch is that day. A schedule fails when nobody entered the window. Its behavior can also be predicted from the config, the property Mercari says it put first [10]. The result Mercari reports is one cluster averaging 44% CPU against a 50% target [11], 6 points under [25]. For that to carry over to another service, its daily curve has to be regular enough to write down as windows, and its campaign start times have to be known before the load arrives [7][9]. TiKV, the storage tier, adds a constraint from TiDB Cloud's control plane. Starting any scale operation puts the cluster into a modifying state, and no further scale operation is accepted until it returns to available [14]. Modifying does not mean down, and Mercari takes care to say the cluster keeps serving [14]. I suspect someone asked. The lock covers the whole cluster, so while a TiKV operation holds it, TiDB cannot scale out even if load rises [15]. Horizontal scale-out adds empty nodes. The cluster returns to available as soon as they join, and rebalancing runs afterwards, outside the lock [16]. Leader moves finish faster than data moves, so leaders appear to shift first [16]. With three copies of the data by default, regions take about three times as long to balance as leaders [17]. Reads go to leaders by default, so read load should even out in about a third of the region-balance time [17]. Scale-in is the slow direction. Every region on the departing node must move off before deletion, and the cluster stays in modifying until the node is gone [18]. Vertical scaling moves no data. It moves leaders off a node and changes the instance class, and is expected to take a few minutes per node [19]. Nodes go one at a time, so a larger cluster holds the lock longer [19]. It can briefly raise load on the remaining nodes and may require PD to scale as well [20]. Its steps are coarse, typically 2x or 4x [23]. Mercari's rule of thumb follows from those costs. Vertical speed depends on node count and not data volume, so it suits a lot of data on few nodes [22]. Horizontal speed depends on data volume and on how many nodes can receive it, so it suits little data on many nodes, where scale-in can beat vertical [22]. Mercari says it will explain its one-method recommendation in a separate post [21].

What to watch

  • Mercari's promised follow-up explaining why TiDB Cloud users should enable only one TiKV scaling method under tidb-operator v1.
  • Any move of TiDB Cloud's scale operations off tidb-operator v1, the basis Mercari gives for its one-method recommendation.
  • Details of Mercari's automated status diagnosis and daily reports, listed among the post's topics.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories