GPT-4: the flagship that raised the bar (2023)
OpenAI's GPT-4 could read images as well as text and scored well on exams built for humans, cementing a fierce commercial race, and a secrecy that unsettled researchers.
Four months after ChatGPT turned language models into a mass phenomenon, OpenAI released GPT-4 on March 14, 2023. It was positioned as a safer, stronger flagship, and it quickly became the model people reached for when they wanted the best available answer.
What was new
GPT-4 pushed on two fronts at once:
- Multimodality. Where earlier large language models handled only text, GPT-4 could also take images as input, describing a photo, reading a chart, or explaining a meme. It signalled that the Transformer recipe was not confined to words.
- Exam-style performance. OpenAI reported that GPT-4 scored in high percentiles on tests designed for humans, from bar exams to advanced placement subjects. Benchmarks like these are imperfect, but the headline numbers reset expectations for what a general model could do.
Under the hood it was still a Transformer scaled up in the tradition of GPT-3, refined with human feedback to behave as a helpful assistant.
The secrecy turn
GPT-4 also marked a shift in how the field shared its work. Citing competition and safety, OpenAI declined to publish the model’s size, architecture, or training data. The accompanying report read more like a product brief than a scientific paper.
For a discipline built on open publication, from the ImageNet benchmark to the Transformer paper, this was a notable change. It fed a wider debate about whether the most capable models would remain inspectable, or become closed commercial products.
Why it matters
GPT-4 confirmed that the assistant unveiled with ChatGPT was not a one-off. It set a capability bar that competitors spent the following year chasing, and its multimodal demos previewed a future where a single model handles text, images, and more.
The releases that followed, from Google’s Gemini to open-weight challengers like Llama 2, are best understood as responses to the standard GPT-4 set.