Carlos Mocholí

@carmocca · User

GitHub profile ↗ · Compare

Research Engineer

Spain1,634 followers33 repositories

Repositories

carmocca/core

:house_with_garden: Open source home automation that puts local control and privacy first.

★ 0Forks 0

carmocca/pytorch

Tensors and Dynamic neural networks in Python with strong GPU acceleration

★ 1PythonForks 0

carmocca/ao

PyTorch native quantization and sparsity for training and inference

★ 0Forks 0

carmocca/neptune-fetcher

Neptune Fetcher is designed to separate data retrieval capabilities from the regular neptune package. This separation makes data fetching more efficient and improves performance.

★ 0Forks 0

carmocca/DeepSpeed

DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

★ 0Forks 0

carmocca/litgpt

Pretrain, finetune, deploy 20+ LLMs on your own data. Uses state-of-the-art techniques: flash attention, FSDP, 4-bit, LoRA, and more.

★ 0Forks 0

carmocca/toolbox

Essential guides and programming tools in my toolbox (with focus on ML Training)

★ 0PythonForks 0

carmocca/litdata

Blazingly fast, distributed streaming of training data from any cloud storage for training AI models

★ 0Forks 0

carmocca/lightning-thunder

Source to source compiler for PyTorch. It makes PyTorch programs faster on single accelerators and distributed.

★ 0Forks 0

carmocca/Fuser

A Fusion Code Generator for NVIDIA GPUs (commonly known as "nvFuser")

★ 0Forks 0

carmocca/TransformerEngine

A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper GPUs, to provide better performance with lower memory utilization in both training and inference.

★ 0Forks 0

carmocca/faster-pytorch-blog

Outlining techniques for improving the training performance of your PyTorch model without compromising its accuracy

★ 0Forks 0

carmocca/DALI

A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.

★ 0Forks 0

carmocca/lightning

Build and train PyTorch models and connect them to the ML lifecycle using Lightning App templates, without handling DIY infrastructure, cost management, scaling, and other headaches.

★ 0PythonForks 0

carmocca/ffcv

FFCV: Fast Forward Computer Vision (and other ML workloads!)

★ 0Forks 0

carmocca/nnutils

CPU & CUDA implementation of several neural network utils

★ 1Forks 1