Skip to content
NS

Noam Shazeer

Thought Leader

Noam Shazeer is a machine learning researcher at OpenAI. He co-authored 'Attention Is All You Need' (2017), the foundational paper introducing the Transformer architecture underlying modern large language models. He co-founded Character.AI, which was acquired in 2024.

Last updated Jul 6, 2026 by the ATDb Editorial Team

Company
Based
San Francisco, California, United States
Connections
2
Years in industry
22 years

Bio

Noam Shazeer is a machine learning researcher at OpenAI. He co-authored 'Attention Is All You Need' (2017), the foundational paper introducing the Transformer architecture underlying modern large language models. He co-founded Character.AI, which was acquired in 2024.

Career

  • Co-founder & CEO

    Character.AI · 2021-2024

  • Senior Research Scientist

    Google Brain / Google · 2000-2021

Expertise & education

Expertise

Large Language Models (LLMs)Transformer ArchitectureMixture of Experts (MoE)Machine Learning InfrastructureAI-Driven PersonalizationNatural Language ProcessingConversational AIAd Relevance and Ranking Systems

Education

  • B.S. Computer Science, Duke University

Speaking topics

Large Language Models and their applicationsTransformer architecture and attention mechanismsEfficient scaling of AI modelsConversational AI and human-computer interaction

Recognition

Notable achievements

  • Co-authored 'Attention Is All You Need' (2017), introducing the Transformer architecture now foundational to all major LLMs
  • Co-founded Character.AI, which was acquired by Google in 2024 for approximately $2.7 billion
  • Contributed to Mixture of Experts (MoE) research enabling efficient large-scale model deployment
  • Key contributor to Google's core machine learning and search ranking infrastructure over two decades

Awards

Test of Time Award — 'Attention Is All You Need' paper widely recognized as one of the most cited and impactful ML papers in history

Publications

  • Vaswani, A., Shazeer, N., et al. — 'Attention Is All You Need' (NeurIPS 2017)
  • Shazeer, N., et al. — 'Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer' (ICLR 2017)
  • Shazeer, N. — 'GLU Variants Improve Transformer' (2020)
Connection details