Benchmark 10 classifiers before trusting one
Sep 2026 ยท 5 min read
Most ML side projects pick a model, report its best number, and move on. For my Diabetes Risk Check I did the boring-honest version: train 10 algorithms plus a 5-expert stacking ensemble, publish the whole comparison table in the repo (train/compare.py โ comparison.json). Losers included.
Why it matters: on a small clinical-style dataset, the gap between algorithms is mostly noise plus preprocessing. Showing all ten keeps you honest about that โ and the ensemble earns its place instead of inheriting it.
Two process details I'm proud of. First, a CTGAN augmentation study tested whether synthetic tabular data actually helps here โ with mixed results that stayed in the write-up instead of being buried. Negative results are results. Second, the trained model is exported to JSON, ported to TypeScript, and a parity gate proves the browser scores exactly what Python scored. A single verify_all.sh (retrain โ export โ parity โ bundle) means the shipped site can never silently drift from the evaluated model.
The whole pipeline reruns from one README block. If a side project can't be reproduced by a stranger, it's a demo, not engineering.
keep reading
Bulk uploads beat dashboards: cutting remittance TAT 80%4 minPending transactions need a tracer4 minCorporate ChatGPT is blocked. So I built ToolChat.7 minI built my own bitly โ with the analytics Bitly hides.7 minGroup chat on REST and polling. No WebSockets. (On purpose.)6 minStreaming LLMs through a 20-line edge proxy.5 min