XGBoost & Gradient Boosting
Boosting vs Bagging
Bagging (Random Forest): trains trees independently in parallel. Reduces variance. Boosting: trains trees sequentially. Each tree corrects the errors of the previous ensemble. Reduces both bias and variance. Gradient Boosting builds trees to fit the negative gradient (residuals) of the loss function.
XGBoost: Extreme Gradient Boosting
XGBoost is the most widely used gradient boosting library. Key innovations: • Second-order Taylor expansion of loss (more precise gradients) • Regularization terms in the objective (L1 and L2 on leaf weights) • Column subsampling (like Random Forest) • Efficient split finding (histogram-based) • Built-in handling of missing values • Parallel tree construction Won hundreds of Kaggle competitions. Go-to for tabular data.
XGBoost vs LightGBM vs CatBoost
XGBoost — Most stable, great docs, wide support. Best default choice. LightGBM — Fastest on large datasets (leaf-wise growth). Memory efficient. CatBoost — Best for categorical features natively. No need to encode.
XGBoost Full Example
early_stopping_rounds halts training when validation loss stops improving — prevents overfitting automatically.
Finished reading? Mark it complete to earn your XP.