Glossary
Short explainers for concepts in the world of artificial intelligence.
A
-
ACM (Association for Computing Machinery)
The main international professional society for computing; awards the Turing Award and publishes journals such as CACM.
-
AI agent
A system that pursues goals by observing an environment, planning, and taking actions (often via tools) rather than only producing static text.
-
AI alignment
The problem of ensuring AI systems pursue intended goals and values, not just what they were literally optimized for on training data.
-
AI winter
A period when AI research funding and public optimism dropped sharply after early hype exceeded practical results.
-
AlexNet
Deep convolutional network that won the 2012 ImageNet challenge and is widely cited as the spark of the modern deep-learning era.
-
AlphaFold
DeepMind system that predicts protein 3D structure from amino-acid sequence; AlphaFold2 (2020) and later releases transformed structural biology.
-
AlphaGo
DeepMind's Go program that beat top professional Lee Sedol in 2016, a landmark mix of deep networks, self-play, and search.
-
Artificial general intelligence (AGI)
Hypothetical AI that matches or exceeds human-level competence across many tasks, not the same as a narrow model or today's LLM chatbot.
-
Artificial intelligence (AI)
The broad field of computer science focused on building systems that perform tasks typically requiring human intelligence, from pattern recognition to reasoning and language understanding.
-
Attention
The mechanism that lets a model weigh how much each part of the input matters to every other part, turning similarity scores into weights via a softmax.
-
Autonomous driving
Vehicle navigation without continuous human steering, combining perception, mapping, planning, and control, often using machine learning on sensor data.
B
-
Backpropagation
Algorithm that trains neural networks by passing the output error backwards through the layers and computing how much each weight contributed to it; the basis of training by gradient descent.
-
Bagging (bootstrap aggregating)
Ensemble technique that trains multiple models on bootstrap samples of the training data and aggregates their predictions to reduce variance.
-
Boltzmann machine
Stochastic neural network with binary units and an energy function from statistical physics; learns hidden representations but typically needs slow Markov-chain (MCMC) sampling to train.
C
-
Chain-of-thought
A prompting and training technique where a model works through intermediate steps in writing before giving a final answer, often improving accuracy on multi-step problems.
-
Chatbot
A conversational interface powered by rules, scripts, or, today often,a large language model; distinct from agents when it only replies without acting in external systems.
-
Convolutional neural network (CNN)
Neural network architecture for grid-like data such as images, using shared local filters and pooling to learn spatial features efficiently.
D
-
DARPA
U.S. Defense Advanced Research Projects Agency; funds high-risk research that has shaped AI, networking, and robotics from the ARPANET era to Grand Challenge races.
-
Dead Internet theory
A fringe hypothesis (from online culture, circa 2021) that most visible web content and engagement is already bots, spam, or AI-generated, not authentic human activity.
-
Decision tree
Supervised model that classifies or predicts by recursively splitting data on feature thresholds, forming a tree of yes-or-no decisions.
-
Deep belief network
Deep neural network built by greedily stacking restricted Boltzmann machines, then fine-tuned with backpropagation; Hinton et al. 2006 helped restart the deep-learning revival.
-
Deep Blue
IBM's 1997 chess supercomputer that beat world champion Garry Kasparov with brute-force search and hand-tuned evaluation, not learning.
-
Deep learning
Neural networks with many successive layers and training methods tuned to stacked representations (often rebranded circa 2006–2012 revival).
-
Deepfake
AI-generated or manipulated image, audio, or video content that resembles real people, objects, places, or events and could falsely appear authentic.
-
DeepMind
London AI lab founded in 2010, acquired by Google in 2014; known for AlphaGo, AlphaFold, and large multimodal models under Google DeepMind.
-
DeepSeek
Chinese AI company whose open-weights reasoning model R1 (2025) rivaled Western frontier systems at a reported fraction of the training cost and shook markets.
E
-
ELIZA
Joseph Weizenbaum's 1966 MIT program that mimicked therapy dialogue via keyword rules, famous for the ELIZA effect.
-
Expert system
Rule-based AI program that encodes domain knowledge as if-then rules and an inference engine; commercial flagship of 1980s symbolic AI.
F
-
Fine-tuning
Additional training starting from an existing model: usually smaller data and narrower objective than pretraining.
G
-
GAN (Generative Adversarial Network)
Two neural networks trained against each other: a generator produces artificial data and a discriminator tries to tell it apart from real data. Introduced by Ian Goodfellow and colleagues in 2014.
-
Generative AI
AI systems that produce new content such as text, images, audio, or video by learning patterns from their training data.
-
GPU
Graphics processing unit; parallel processor originally built for rendering images, later repurposed to accelerate matrix operations in neural network training.
-
Gradient descent
Optimisation method that repeatedly moves a model's parameters a small step in the direction that lowers the error most; the step size is set by the learning rate.
-
Guardrails
Technical and policy layers that constrain what an AI system may say or do: filters, rules, tool permissions, and human review.
H
-
Hallucination
In AI, fluent but false or fabricated model output, not sensory hallucination; common risk for LLMs, chatbots, and cited answers.
I
-
IBM Watson
IBM's question-answering system that beat Jeopardy! champions Ken Jennings and Brad Rutter in 2011; IBM later used the name for its commercial AI offerings.
-
ImageNet
Large labelled image dataset and benchmark (mid-2000s onward) that made progress in computer vision measurable and paved the way for deep learning.
-
Inference
Using a trained model to produce outputs on new inputs: after training is finished (also called deployment or forward pass in many setups).
J
-
Jailbreak
Crafting prompts (or prompt chains) intended to bypass a model's safety rules, refusals, or usage policies, not the same as ordinary prompt engineering.
K
-
Kernel trick
Technique that computes inner products in a high-dimensional feature space implicitly, enabling linear algorithms to fit nonlinear decision boundaries efficiently.
L
-
Language model
Statistical or neural model that assigns probabilities to word sequences, used for prediction, generation, and as a backbone for modern NLP systems.
-
Large language model (LLM)
A statistical language model trained on large text corpora, typically using a transformer architecture, used for generation and understanding tasks.
-
LeNet
Family of convolutional neural networks by Yann LeCun and colleagues at AT&T Bell Labs; LeNet-1 read zip-code digits around 1989, LeNet-5 read bank cheques in 1998.
-
LIDAR
Ranging sensor that measures distance by timing laser pulses, producing dense 3D point clouds used in mapping, robotics, and autonomous vehicles.
M
-
Machine learning (ML)
Algorithms that generalize patterns from data to make predictions or decisions without hand-written rules for every case.
-
Markov chain
A random process where the next state depends only on the current state; in machine learning, the basis of MCMC sampling, early n-gram language models, and slow iterative generators that GANs were designed to avoid.
-
Mixture of experts
A model design that splits the network into many specialized subnetworks and activates only a few per input, so a very large model can run at the cost of a much smaller one.
-
Model (machine learning model)
A learned function with parameters fit on data: the artifact produced by training and used at inference time.
-
Model Context Protocol (MCP)
An open protocol that standardises how AI applications connect to external data sources, tools, and services, enabling models to act on live context rather than static training data.
-
Model router
A control layer that picks which model variant or mode to run for each request, for example fast chat versus deeper reasoning, instead of leaving the choice to the user.
N
-
Natural language processing (NLP)
The branch of AI focused on understanding, generating, and transforming human language with computers, from parsing and translation to modern LLMs.
-
Neural network (ANN)
Parametric models organized in layers of simple units; strengths are learned from data via optimization.
P
-
Perceptron
Early single-layer neural model that learns a linear decision boundary from examples; Rosenblatt (1958), later limited by Minsky and Papert's analysis.
-
PPO (Proximal Policy Optimization)
A widely used reinforcement-learning algorithm that stabilizes policy updates with a clipped objective; common in RLHF fine-tuning stages.
-
Prompt
The textual instruction or context given to an AI model to guide output style, scope, and task.
-
Prompt injection
Hostile or hidden instructions in context that hijack an LLM's behavior, direct (user input) or indirect (RAG, tools, untrusted documents).
Q
-
Question answering
Area of natural language processing concerned with systems that answer questions asked in natural language with a direct answer rather than a list of documents.
R
-
Random forest
Ensemble classifier that aggregates many decorrelated decision trees trained on bootstrap samples with random feature subsets at each split.
-
Reasoning
Step-by-step inference behavior used to derive conclusions from context, rules, or intermediate states.
-
Red teaming
Deliberately probing an AI system for failures, misuse, and bypasses before or after launch, structured adversarial testing, not casual chatting.
-
Reinforcement learning (RL)
Learning policies via rewards or scores from interacting with an environment: core to games, robotics, and some RLHF-style LLM tuning.
-
ReLU (rectified linear unit)
Activation f(x) = max(0, x) widely used in deep nets; popularized mid-2010s as a fast alternative to sigmoid/tanh.
-
Retrieval-augmented generation (RAG)
Combining a retriever over documents or tools with a generator LLM so answers can cite fresher or private context.
-
RLHF (reinforcement learning from human feedback)
Fine-tuning LLMs from human preference comparisons, often followed by reinforcement learning, widely used for alignment after pretraining.
S
-
Self-play
Training regime in which an agent improves by competing against copies or past versions of itself, generating diverse experience without human expert labels.
-
Sovereign AI
National or locally controlled AI capacity, including infrastructure, data, models, and policy, built to serve a country's interests rather than depend entirely on foreign stacks.
-
Superintelligence (ASI)
A hypothetical AI that far exceeds human capability across virtually all cognitive domains, not the same as a strong chatbot, today's LLM, or human-level AGI.
-
Supervised learning
Machine learning from labeled input–output pairs: the model learns to map examples to known targets before deployment.
-
Support vector machine (SVM)
Supervised classifier that finds a maximum-margin boundary between classes, often extended to nonlinear problems via the kernel trick.
-
Symbolic AI (good old-fashioned AI)
AI built from explicit symbols, logic rules, and structured knowledge: often contrasted with numeric learning from raw data.
T
-
Technological singularity
A hypothetical future point where technological change (often via advanced AI) accelerates beyond reliable human forecasting or control.
-
Temporal-difference learning (TD learning)
Reinforcement-learning method that updates value estimates from the difference between successive predictions, bootstrapping from experience without waiting for a final reward.
-
Test-time compute
Extra computation spent when a model answers a question, at inference time, rather than only during training; a lever for harder reasoning tasks.
-
Token
The smallest text unit a language model processes (often a word piece, not a full word).
-
Training (model training)
The phase where model parameters are optimized on data or feedback before deployment: distinct from inference at run time.
-
Transformer (architecture)
Neural sequence model built mainly on attention: parallelizable and foundational for LLMs since "Attention Is All You Need" (2017).
-
Turing test (imitation game)
Alan Turing’s proposed behavioral criterion: can dialogue from a machine be distinguished from a human’s?
U
-
Unrolled inference network
An iterative inference procedure in an energy-based or probabilistic model, unrolled into a fixed number of steps and trained end to end as a feed-forward network; GANs avoid this by generating in one forward pass.
-
Unsupervised learning
Machine learning from unlabeled data: the model finds structure such as clusters or patterns without being told the right answers.
V
-
Vector database
A store optimized for similarity search over embedding vectors: the retrieval layer behind many RAG and semantic-search systems.
W
-
Word embedding
Dense vector representation of a word learned from co-occurrence or prediction context, capturing semantic and syntactic regularities in continuous space.