Skip to content
NS

Noam Shazeer

Thought Leader

Noam Shazeer is a machine-learning researcher and co-author of 'Attention Is All You Need', the paper that introduced the Transformer. He co-founded Character.AI, returned to Google as a VP of Engineering and co-lead of Gemini, and in June 2026 left Google to join OpenAI.

Last updated Sep 26, 2026 by the ATDb Editorial Team

Company
Based
San Francisco, California, United States
Connections
2
Years in industry
22 years

Bio

Noam Shazeer is a machine-learning researcher and co-author of 'Attention Is All You Need', the paper that introduced the Transformer. He co-founded Character.AI, returned to Google as a VP of Engineering and co-lead of Gemini, and in June 2026 left Google to join OpenAI.

Career

  • Co-founder & CEO

    Character.AI · 2021-2024

  • Senior Research Scientist

    Google Brain / Google · 2000-2021

Expertise & education

Expertise

Large Language Models (LLMs)Transformer ArchitectureMixture of Experts (MoE)Machine Learning InfrastructureAI-Driven PersonalizationNatural Language ProcessingConversational AIAd Relevance and Ranking Systems

Education

  • B.S. Computer Science, Duke University

Speaking topics

Large Language Models and their applicationsTransformer architecture and attention mechanismsEfficient scaling of AI modelsConversational AI and human-computer interaction

Recognition

Notable achievements

  • Co-authored 'Attention Is All You Need' (2017), introducing the Transformer architecture now foundational to all major LLMs
  • Co-founded Character.AI, which was acquired by Google in 2024 for approximately $2.7 billion
  • Contributed to Mixture of Experts (MoE) research enabling efficient large-scale model deployment
  • Key contributor to Google's core machine learning and search ranking infrastructure over two decades

Awards

Test of Time Award — 'Attention Is All You Need' paper widely recognized as one of the most cited and impactful ML papers in history

Publications

  • Vaswani, A., Shazeer, N., et al. — 'Attention Is All You Need' (NeurIPS 2017)
  • Shazeer, N., et al. — 'Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer' (ICLR 2017)
  • Shazeer, N. — 'GLU Variants Improve Transformer' (2020)
Connection details