Gemini 1.0: Google's multimodal answer (2023)

Google rebranded its AI efforts under Gemini, a natively multimodal family meant to rival GPT-4, and learned how closely launch demos would now be scrutinised.

On December 6, 2023, Google DeepMind announced Gemini 1.0, the company’s most direct answer to GPT-4. It was presented as a single family in three sizes, from a compact version meant to run on phones up to a top-end model aimed squarely at the frontier.

Built multimodal from the start

Google’s pitch was that Gemini was natively multimodal: rather than bolting image handling onto a text model, it was designed from the beginning to work across text, images, audio, and more. In practice this is a difference of degree more than kind, since GPT-4 was already multimodal, but it signalled where the large language model race was heading, toward assistants that treat different kinds of input as one.

The launch also folded Google’s scattered AI branding into one name. The consumer assistant previously called Bard was moved under the Gemini label, and the models were slated to appear across Google’s products, from search to Android phones.

The demo controversy

The announcement leaned on a slick video in which Gemini appeared to react, in real time, to drawings and objects. Reporting soon revealed the demo had been edited: prompts were partly text, and responses were sped up and trimmed. Google acknowledged the video was not a real-time capture.

The episode became a small lesson of the era. With so much money and attention riding on each release, launch materials were now examined as closely as the models, and the gap between a polished demo and everyday behaviour became a story in itself.

Why it matters

Gemini confirmed that the Transformer-driven assistant race was now a contest among the largest technology companies, not a single lab’s project. Google had the data, the distribution, and the hardware to compete, and Gemini was the vehicle.

It also marked the moment multimodality became table stakes rather than a headline feature. From here on, a serious frontier model was expected to see and hear, not just read.

#llm #multimodal