Skip to content
Piyush ๐Ÿ‘‹
all writing

Benchmark 10 classifiers before trusting one

Sep 2026 ยท 5 min read

Most ML side projects pick a model, report its best number, and move on. For my Diabetes Risk Check I did the boring-honest version: train 10 algorithms plus a 5-expert stacking ensemble, publish the whole comparison table in the repo (train/compare.py โ†’ comparison.json). Losers included.

Why it matters: on a small clinical-style dataset, the gap between algorithms is mostly noise plus preprocessing. Showing all ten keeps you honest about that โ€” and the ensemble earns its place instead of inheriting it.

Two process details I'm proud of. First, a CTGAN augmentation study tested whether synthetic tabular data actually helps here โ€” with mixed results that stayed in the write-up instead of being buried. Negative results are results. Second, the trained model is exported to JSON, ported to TypeScript, and a parity gate proves the browser scores exactly what Python scored. A single verify_all.sh (retrain โ†’ export โ†’ parity โ†’ bundle) means the shipped site can never silently drift from the evaluated model.

The whole pipeline reruns from one README block. If a side project can't be reproduced by a stranger, it's a demo, not engineering.