Machine learning (ML) is a field of study in artificial intelligence concerned with the development and study of statistical algorithms that can learn from data and generalize to unseen data, and thus perform tasks without being explicitly programmed.[1] Advances in the field of deep learning have allowed neural networks, a class of statistical algorithms, to surpass many previous machine learning approaches in performance.
Statistics and mathematical optimisation methods compose the foundations of machine learning. Data mining is a related field of study, focusing on exploratory data analysis (EDA) through unsupervised learning.[3][4]
From a theoretical viewpoint, probably approximately correct learning provides a mathematical and statistical framework for describing machine learning. Most traditional machine learning and deep learning algorithms can be described as empirical risk minimisation under this framework.
History
For a chronological guide, see Timeline of machine learning.
The term machine learning was coined in 1959 by Arthur Samuel, an IBM employee and pioneer in the field of computer gaming and artificial intelligence.[5][6] The synonym self-teaching computers was also used during this time period.[7][8]
The earliest machine learning program was introduced in the 1950s, when Samuel invented a computer program that calculated the chance of winning in checkers for each side, but the history of machine learning is rooted in decades of efforts to study human cognitive processes.[9] In 1949, Canadian psychologist Donald Hebb published the book The Organization of Behavior, in which he introduced a theoretical neural structure formed by certain interactions among nerve cells.[10] The Hebbian theory of neuron interaction set the groundwork for how many machine learning algorithms work, with connected artificial neurons changing the strength of their connections based on data.[9] Other researchers who have studied human cognitive systems contributed to the modern machine learning technologies as well, including Walter Pitts and Warren McCulloch, who proposed the first mathematical model of neural networks including algorithms that mirror human thought processes.[9][verification needed]
By the early 1960s, an experimental “learning machine” with punched tape memory, called Cybertron, had been developed by Raytheon Company to analyse sonar signals, electrocardiograms, and speech patterns using rudimentary reinforcement learning. It was repetitively “trained” by a human operator/teacher to recognise patterns and equipped with a “goof” button to cause it to reevaluate incorrect decisions.[11][importance?] A representative book[citation needed] on research into machine learning during the 1960s was Nils Nilsson‘s book “Learning Machines”, dealing mostly with machine learning for pattern classification.[12] Interest related to pattern recognition continued into the 1970s, as described by Duda and Hart in 1973.[13] In 1981, a report was given on using teaching strategies so that an artificial neural network learns to recognise 40 characters (26 letters, 10 digits, and 4 special symbols) from a computer terminal.[14]
Tom M. Mitchell provided a widely quoted,[citation needed] more formal definition of the algorithms studied in the machine learning field: “A computer program is said to learn from experience E with respect to some class of tasks T and performance measure P if its performance at tasks in T, as measured by P, improves with experience E.”[15] This definition of the tasks in which machine learning is concerned is fundamentally operational rather than defining the field in cognitive terms. This follows Alan Turing‘s proposal in his paper “Computing Machinery and Intelligence“, in which the question, “Can machines think?”, is replaced by asking whether machines can convincingly imitate a human in its responses to human-posed questions.[16][17]
In 2012, AlexNet, developed by Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton, achieved substantially improved results in the ImageNet image recognition competition, contributing to the wider adoption of deep neural networks.[18] In 2013, Tomáš Mikolov and colleagues introduced word2vec, techniques for efficiently learning distributed vector representations of words from large text corpora.[19][20] In 2014, Ian Goodfellow and colleagues introduced generative adversarial networks (GANs), a framework for training generative models through an adversarial process.[21] In 2016, AlphaGo became the first computer program to defeat a professional human Go player without handicaps on a full-sized board, using deep neural networks and reinforcement learning.[22] In 2017, Ashish Vaswani and colleagues introduced the Transformer, a neural network architecture based primarily on attention rather than recurrence or convolution.[23]
