Transformer is a neural network architecture introduced in 2017 that, using a mechanism known as “attention,” can effectively process long sequences of text and take into account relationships even between distant words in a sentence or document. This architecture forms the basis of nearly all modern large language models, including GPT, Claude, and Gemini, and its invention is considered one of the key milestones in the current boom in generative AI.