zhongbozhu/Megatron-LM
Ongoing research training transformer models at scale
DevTech @ NVIDIA | Ex Amazon | MS CompE @ UIUC | ZJUer
Ongoing research training transformer models at scale
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper and Ada GPUs, to provide better performance with lower memory utilization in both training and inference.
Training library for Megatron-based models with bidirectional Hugging Face conversion capability
Mixture-of-experts (MoE) training megakernel for NVL72s
Scalable toolkit for efficient model reinforcement
Megatron's multi-modal data loader
DeepEP: an efficient expert-parallel communication library
Online Course GAMES101 Assignment7: Path Tracing Algorithm Implemented with Multi-thread Acceleration
HarmoniOS - tiny Linux-like OS kernel
Simulate TCP reliable data transfer with UDP
mmap for file access && mmap for IPC
Apache Hadoop
Raft consensus implementation using Java, tested in framework based on python asyncio
learning inter-process communication using pipes between nodejs and python
deploy the quora question pair model to web app
code for wrap fputs in C to a python module using python.h and distutils
This repository contains the solutions and explanations to the algorithm problems on LeetCode. Only medium or above are included. All are written in C++/Python and implemented by myself. The problems attempted multiple times are labelled with hyperlinks.