Support vector machines: maximum-margin classifiers with kernels (1995)
In 1995 Corinna Cortes and Vladimir Vapnik published support-vector networks in Machine Learning, combining maximum-margin linear separation with the kernel trick for nonlinear boundaries.
In 1995, Corinna Cortes and Vladimir Vapnik published Support-Vector Networks in Machine Learning (volume 20, issue 3). The paper presented support vector machines (SVMs): classifiers that find a maximum-margin boundary between categories and represent data only through the training points that lie on that margin, the support vectors. The method built on earlier work by Bernhard Boser, Isabelle Guyon, and Vapnik on optimal margin classifiers (1992), which introduced the kernel trick to handle nonlinear decision surfaces efficiently.
Margin, kernels, and sparse solutions
Linear classifiers can fail when classes are not separable by a straight line in the input space. SVMs map inputs into a higher-dimensional feature space via a kernel function so that a linear separator in that space corresponds to a curved boundary in the original space. The kernel trick computes inner products in the expanded space without building the full mapping explicitly. The resulting model depends on a subset of training examples, which made SVMs attractive when data were scarce and neural networks were out of fashion during the second AI winter (article).
A statistical rival to neural nets
Through the late 1990s and 2000s, SVMs became a standard tool in machine learning for text classification, bioinformatics, and vision benchmarks, often competing with random forests (article) and later with deep learning. They exemplified the shift from hand-crafted expert systems toward supervised learning on labelled data.
Why it matters
Support vector machines brought rigorous statistical learning theory into widespread practice. Vapnik’s earlier work on structural risk minimisation helped justify why maximising margin could generalise well. Even after AlexNet revived neural networks, SVM ideas influenced how researchers thought about margins, kernels, and high-dimensional classification.