A gradient-boosted ensemble that turned an overlooked feature — sky position — into the largest single source of improvement, then layered pseudo-labeling and a genuinely diverse third model on top.
4,500 sampled training points by sky position (α right ascension, δ declination). Dense, class-pure regions are visible by eye — removing this signal cost 1.4 points of balanced accuracy.
Each row is a fully cross-validated 5-fold balanced accuracy score on real labels — pseudo-labeled rows, where used, were always held out of validation.
Confusion matrix for the final CatBoost + LightGBM-v3 blend, out-of-fold on original labels. Rows are true class, columns are predicted class.
| GALAXY | QSO | STAR | |
|---|---|---|---|
| GALAXY | 363,83396.4% | 5,0191.3% | 8,6282.3% |
| QSO | 1,9041.6% | 114,18597.5% | 1,0540.9% |
| STAR | 2,3982.9% | 3830.5% | 79,94396.6% |
Balanced accuracy is the average of these three numbers — STAR is the hardest class, most often confused with GALAXY at the boundary.
Gain importance, normalized to the top feature. Cyan bars are the spatial KNN features this project added.
Stars sit near zero redshift; quasars skew high; galaxies span the middle — with enough overlap that redshift alone can't fully separate the classes.
Inverse-frequency class weighting during training kept the model from defaulting to the majority class — predicted proportions track the training distribution closely.