RuiWang1998/Deep-Approximate-Shapley-Propagation
This is a Pytorch Implementation of the DASP algorithm from the paper "Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Value Approximation"
This is a Pytorch Implementation of the DASP algorithm from the paper "Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Value Approximation"
Fast and memory-efficient exact attention
Ongoing research training transformer models at scale
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper and Ada GPUs, to provide better performance with lower memory utilization in both training and inference.
CUDA kernels for the Longhorn Architecture
Official PyTorch Implementation of the Longhorn Deep State Space Model
A PyTorch Extension: Tools for easy mixed precision and distributed training in Pytorch
Central place for the engineering/scaling WG: documentation, SLURM scripts and logs, compute environment and data.
OmegaFold Release Code
Google Research forked
A set of examples around pytorch in Vision, Text, Reinforcement Learning, etc.