Kaggle Playground Series S6E6

Classifying galaxies, quasars,
and stars from sky survey data

A gradient-boosted ensemble that turned an overlooked feature — sky position — into the largest single source of improvement, then layered pseudo-labeling and a genuinely diverse third model on top.

Balanced Accuracy
0.9683
↑ +0.00435 vs. baseline

The discovery that drove this project

4,500 sampled training points by sky position (α right ascension, δ declination). Dense, class-pure regions are visible by eye — removing this signal cost 1.4 points of balanced accuracy.

α (right ascension) → ↑ δ (declination)
The journey

Six steps from baseline to final blend

Each row is a fully cross-validated 5-fold balanced accuracy score on real labels — pseudo-labeled rows, where used, were always held out of validation.

00
LightGBM baseline
Color indices + redshift transforms only
0.96398
01
+ spatial cartesian + KNN density
Sky position turned out to carry a ~1.4pt signal
0.96681 +0.00283
02
XGBoost, same features
Comparison base model
0.96643 +0.00245
03
+ pseudo-labeling
194K test rows added at ≥99% confidence
0.96707 +0.00309
04
+ CatBoost (diverse model)
Ordered boosting -- only 97.9% agreement with LightGBM
0.96781 +0.00383
05
Final blend: CatBoost + LightGBM-v3
Simple average beat a tuned weight search on held-out check
0.96833 +0.00435
Model comparison

Every model trained

LightGBM v1 (baseline)
0.96398
XGBoost
0.96643
LightGBM v2
0.96681
LightGBM v3 (pseudo)
0.96707
CatBoost (pseudo)
0.96781
Final Blend
0.96833
Why CatBoost was worth the extra hour

Diversity, not just averaging

LightGBM and XGBoost, trained on identical features and folds, agreed on 99.4% of predictions — almost nothing for an ensemble to exploit. A stacked meta-learner on top of them actually hurt the score.

CatBoost's ordered boosting agreed with LightGBM on only 97.9% of predictions. It became both the strongest standalone model and the model that improved the blend the most.
Final blend performance

Where the errors still live

Confusion matrix for the final CatBoost + LightGBM-v3 blend, out-of-fold on original labels. Rows are true class, columns are predicted class.

Confusion matrix

GALAXYQSOSTAR
GALAXY
363,83396.4%
5,0191.3%
8,6282.3%
QSO
1,9041.6%
114,18597.5%
1,0540.9%
STAR
2,3982.9%
3830.5%
79,94396.6%

Per-class recall

Balanced accuracy is the average of these three numbers — STAR is the hardest class, most often confused with GALAXY at the boundary.

GALAXY recall 0.9638
QSO recall 0.9747
STAR recall 0.9664
Feature importance

What the model actually uses

Gain importance, normalized to the top feature. Cyan bars are the spatial KNN features this project added.

redshift
100.0
redshift_log
88.8
g_i
77.2
redshift_sqrt
34.2
u_r
18.8
band_std
13.5
knn25_STAR
10.2
knn25_GALAXY
9.9
u_z
9.0
z
7.5
g
7.2
redshift_x_gr
5.8
Redshift by class

The strongest single feature

Stars sit near zero redshift; quasars skew high; galaxies span the middle — with enough overlap that redshift alone can't fully separate the classes.

Class balance

Training distribution vs. final predictions

Inverse-frequency class weighting during training kept the model from defaulting to the majority class — predicted proportions track the training distribution closely.

Training set
65%
20%
14%
Test predictions (final blend)
64%
21%
16%