Loading timeline…
20202020s
Alexey Dosovitskiy
Lead author of the Vision Transformer, which showed images could be handled by the architecture built for language.
Organizations
Google BrainIntel LabsInceptive
Major Achievements
- •Led 'An Image is Worth 16x16 Words' (2020), applying a pure Transformer to image patches.
- •Demonstrated that with enough data, Transformers beat convolutional networks at image recognition.
- •The result began the convergence of vision and language onto a single architecture.
- •Earlier developed FlowNet, bringing deep learning to optical flow estimation.
Key Papers
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy · ICLR 2021 · 2020