Build1 publisher3 min readPublished
Guessing 'no' for every patient scored 95.3% accuracy on public hypothyroid data
Default boosted trees came within half an AUC point of a 200-fit hyperparameter search on five of six public datasets, a dev.to benchmark found. Anyone reviewing an AI-generated modelling notebook can check the cheap baselines and the default trees first, and treat the tuning search as a cost it has to justify.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Gradient-boosted trees on scikit-learn defaults beat an untuned logistic regression by 2 to 15 AUC points across the six datasets.
- On telecom churn, predicting that nobody leaves scored 85.9% accuracy and the logistic regression scored 86.6%.
- The default trees fit in under a second on each dataset, while the tuning search took 43 to 152 seconds per dataset on a 4-core machine.
- According to the post, an AI coding assistant asked to build a churn model often returns a finished-looking notebook with boosted trees, a tuning search and a classification report.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint On problems this imbalanced, accuracy cannot be the headline metric in a model report: a churn model presented as 87% accurate beats predicting that nobody leaves by only 0.7 points.
- cost On every dataset the search used at least 43 times the wall-clock of a default fit, and on five of six it bought an average gain of -0.3 to +0.2 AUC points for that time.
- exposure A generated notebook with no feature-ignoring baseline can clear review on an accuracy figure that sits within a point of doing nothing.
On an imbalanced dataset, accuracy mostly measures the class balance. Positive cases are 4.8% of the hypothyroid data [2]. A model that answers "no" for every patient is 95.3% accurate [1]. The untuned logistic regression reached 96.2% [3], 0.9 points above a model that never reads a feature [1]. AUC works differently. The majority guess scores 0.500 by definition, because it gives every row the same score [13].
scikit-learn ships that floor as DummyClassifier. Its documentation says the class "makes predictions that ignore the input features" and "serves as a simple baseline to compare against other more complex classifiers" [9]. The post's author wants that model found before any number in a generated notebook gets read. "If there isn't one, none of the numbers mean anything yet," the author wrote [21]. The author does not blame the tool. According to the author, the assistant did nothing wrong: it answered the question it was asked, and a request to build a model does not ask for a floor to measure it against [15].
The default rung needs a look at the defaults. On the two datasets with more than 10,000 rows, HistGradientBoostingClassifier turns on early stopping without being asked [10]. On those two, "default" already includes a small amount of built-in tuning [10]. I think some of the search's near-zero gain is credit to scikit-learn's defaults, which left the 200 fits little to find [6].
The test harness has details a reviewer should look for in any notebook. Each dataset got one stratified 80/20 split. The search saw only the training data, and every score comes from the held-out 20% [11]. The churn data's phone-number column, a unique ID per customer, was dropped [12]. Categorical columns were one-hot encoded for the logistic regression and passed to the trees as native categories [12]. The search-gain figure was checked across four random splits, beyond the single split in the table [6].
This is someone else's workload. The six datasets are public binary-classification sets from the Penn Machine Learning Benchmark collection [16], run on scikit-learn 1.9.1 [17]. For the tuning result to transfer, a problem has to share three conditions: it is tabular, it is binary, and default trees already fit it well. I would rerun the comparison before trusting the search result on anything else. The baseline check applies to any classifier, since the guess and the logistic regression together take seconds [18].
The author describes the four models as a ladder. "Each rung of that ladder answers a different question, and the rung below it supplies the number you need to answer it," the author wrote [20]. In review order:
1. The majority-class guess, found or fitted first. 2. The logistic regression, measured against that floor. 3. The default trees, measured against the logistic regression. 4. The search, which has to justify 40 configurations times five folds against the default trees [7].
The author calls the search "the step that makes a notebook look thorough" [19]. On five of these six datasets, looking thorough was most of what it added [6].
What to watch
- Per-dataset AUC figures for the jump from logistic regression to default trees, which would show which datasets sit at the 15-point end of the range.
- A rerun on larger or messier tabular data where default trees fit poorly, which would test whether the search stays near zero outside these six datasets.
- Whether AI coding assistants start adding a DummyClassifier baseline to generated modelling notebooks without being asked.