Bowen12992/flash-attention
Fast and memory-efficient exact attention
Fast and memory-efficient exact attention
FlagScale is a large model toolkit based on open-sourced projects.
FlagGems is an operator library for large language models implemented in Triton Language.
The Runner for GitHub Actions :rocket:
This is the source code for the paper entitled "Span-based Named Entity Recognition by Generating and Compressing Information"
some test with pytorch
Development repository for the Triton language and compiler
上海交通大学 XeLaTeX 学位论文及课程论文模板 | Shanghai Jiao Tong University XeLaTeX Thesis Template
Depth of field simulator
Open deep learning compiler stack for cpu, gpu and specialized accelerators
Super-project for modularized Boost
:books: 华章计算机科学丛书高清扫描
Tensors and Dynamic neural networks in Python with strong GPU acceleration
The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.
PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)
CUDA Templates for Linear Algebra Subroutines