Scaling fraud model experimentation 4× at Shopify

May 3 20268 min read

At Shopify, we were running hyperparameter optimization for one of our LightGBM fraud models sequentially. A typical study would get through around 30 trials in 22 hours.

The problem wasn't just that it was slow. The improvements we were getting from those runs were small enough that, when iterating on the model, rerunning HPO often wasn't worth losing almost a full day. In practice, that meant we were doing less hyperparameter search than we probably should have.

I spent some time trying to fix that.

I rebuilt the job so that a single Optuna study could run across four Vertex AI machines, with each machine contributing trials to the same optimization history. That took us from roughly 30 trials to around 120 in the same 22-hour window.

Then the model barely improved.

It turned out compute wasn't the only bottleneck. Running four times as many trials had mostly just searched the same constrained parameter space more thoroughly. The distributed system made experimentation much faster, but it also exposed a second problem: the search space itself.

One Bayesian search

This was a Bayesian optimization loop. Optuna used the results of completed trials to decide which regions of the search space looked promising. Every new score changed what it tried next.

That dependency shaped the system. Four separate studies would learn from four partial histories. I needed four machines contributing to one optimizer.

1

Sample

Bayesian sampler proposes a configuration

2

Train

LightGBM fits on the training data

3

Evaluate

Score the objective on held-out data

4

Update

Commit the trial to shared history

Completed objectives inform what the sampler proposes next.

Distributing the loop

Each Vertex node read its role from the cluster spec and joined the same Optuna study. Optuna's Cloud SQL-backed RDB storage held the global trial state and objective history. A node claimed a trial, trained and scored its LightGBM model, then committed the result for every other node to learn from.

Vertex AI worker pool

Chief

trial A · train + score

Worker 1

trial B · train + score

Worker 2

trial C · train + score

Worker 3

trial D · train + score

pull next trial · write result

Shared Optuna study

suggest · claim · record

read / write history

Cloud SQL

shared optimization history

Trial budget reached
Chief retrains the best configuration
Model + study artifacts
Each free node pulls its next trial from the same Cloud SQL history.

The workers ran asynchronously. Some configurations were pruned early while others trained much longer. Splitting the budget into four fixed chunks would leave faster nodes idle behind stragglers.

I used pull-based dynamic work stealing instead. A node asked the shared study for a new trial as soon as it became free. Work followed available compute, so the pool kept moving without a central dispatcher assigning fixed batches.

That pull model created a global stopping race. Near the end of the study, several workers could see remaining budget and start a trial at the same time. I added overshoot guards around the shared trial count so the asynchronous loop stopped at the intended budget.

The chief participated in the search like any other node. Once the budget closed, the workers exited and the chief loaded the best completed trial. It retrained that configuration on the full dataset and uploaded the final model and study artifacts to GCS.

Sequential
~30 trials
Four nodes
~120 trials
Roughly the same day-long training window

The search hit a wall

I had run 120 trials over the same ranges used by the old 30-trial search. I expected the extra attempts to find a better configuration. They did not.

Optuna's slice plots explained why. Each dot was one trial. The horizontal position showed the value Optuna tried, and the vertical position showed its score. Many of the strongest dots were hugging the right side of the plot. Optuna kept finding better results near the largest value it was allowed to test.

That is a useful failure mode. The optimizer was telling us that the interesting part of the search probably continued past the boundary.

trialstrongest trials

Old ranges

The strongest trials stack up at the right boundary.

before

parameter 1

parameter 2

parameter 3

Relative importance

Wider ranges

The strongest trials settle inside the range.

after

parameter 1

parameter 2

parameter 3

Relative importance

Conceptual recreation based on the talk. Parameter names, ranges, and values are omitted.

Opening up the search

I widened the constrained ranges and ran the study again. This time, the strongest trials moved away from the edge and settled inside the allowed space. Parameter importance also became less concentrated, which was another sign that the optimizer had room to explore.

There was nothing magic about four nodes or 120 trials. They were practical starting points. The useful change was getting the feedback loop down to a reasonable cadence: run a deeper search, inspect its shape, adjust the space, and run it again.

Scoring the result

Accuracy is a bad metric for fraud because most transactions are legitimate. I used precision-recall area under the curve, or PR-AUC, to measure how well the model found fraud while controlling false positives across thresholds.

I also tracked a GMV-weighted version. A missed $20 fraud attempt and a missed $2,000 attempt count equally in regular PR-AUC. Weighting by gross merchandise value adds a view of the financial exposure.

4×trial throughput
+8.8%mean PR-AUC
+33.8%mean GMV-weighted PR-AUC

After widening the search space, the final study improved mean PR-AUC by 8.8% and mean GMV-weighted PR-AUC by 33.8%.

My first guess was that more trials would automatically produce a better model. I was wrong. More trials made the bad assumption easier to see. The four-node system gave us enough speed to find that mistake, fix it, and search again while the result was still useful.