Mesh TensorFlow: Model Parallelism Made Easier
Repositories
lucidrains repositories
Implementation of MeshGPT, SOTA Mesh generation using Attention, in Pytorch
Implementation of the MetaController proposed in "Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning" from the Paradigms of Intelligence team at Google
Implementation of Metaformer, but in an autoregressive manner
Implementation of MetNet-3, SOTA neural weather model out of Google Deepmind, in Pytorch
Implementation of Mimic-Video, Video-Action Models for SOTA Generalizable Robot Control Beyond VLAs
Implementation of the proposed minGRU in Pytorch
Implementation of Mind Evolution, Evolving Deeper LLM Thinking, from Deepmind
Implementation of 🌻 Mirasol, SOTA Multimodal Autoregressive model out of Google Deepmind, in Pytorch
Some personal experiments around routing tokens to different autoregressive attention, akin to mixture-of-experts
A Pytorch implementation of Sparsely-Gated Mixture of Experts, for massively increasing the parameter count of language models
An implementation of masked language modeling for Pytorch, made as concise and simple as possible
A GPT, made only of MLPs, in Jax
An All-MLP solution for Vision, from Google AI
Implementation of a single layer of the MMDiT, proposed in Stable Diffusion 3, in Pytorch
Usable implementation of Mogrifier, a circuit for enhancing LSTMs and potentially other networks, from Deepmind
Pytorch reimplementation of Molecule Attention Transformer, which uses a transformer to tackle the graph-like structure of molecules
Multi-Joint dynamics with Contact. A general purpose physics simulator.
Implementation of a multimodal diffusion transformer in Pytorch
Implementation of Multiscreen proposed by Ken Nakanishi for "Screening is Enough"