Random forests: Leo Breiman's ensemble of decision trees (2001)
In October 2001 Leo Breiman published Random Forests in Machine Learning, combining many decorrelated decision trees trained on bootstrap samples to create a powerful general-purpose classifier.
In October 2001, Leo Breiman published Random Forests in Machine Learning (volume 45, issue 1). The method builds a large ensemble of decision trees, each trained on a random bagging sample of the data and considering only a random subset of features at each split. Predictions are combined by majority vote or averaging, which reduces variance compared with a single tree.
Why randomness helps
A lone decision tree can overfit noise by growing deep branches that memorise training quirks. Random forests deliberately inject randomness so individual trees make different errors. Aggregating many decorrelated trees often yields robust accuracy on tabular data without the feature engineering demanded by early expert systems. The approach sits firmly in supervised learning and classical machine learning, parallel to support vector machines (article).
Before the deep-learning wave
Random forests became a workhorse algorithm in the 2000s for bioinformatics, remote sensing, and business analytics, years before ImageNet and AlexNet made deep learning dominant in vision. They remain common when datasets are modest in size and interpretability of feature importance matters.
Why it matters
Breiman’s paper unified earlier ideas about bagging and random feature subsets into a simple, strong default method. It showed that ensemble averaging could match or beat single complex models, a theme that later reappeared in other domains. The publication date in 2001 places it in the quiet statistical revival between the second AI winter and the deep belief network era (article).