carmocca/core
:house_with_garden: Open source home automation that puts local control and privacy first.
Research Engineer
:house_with_garden: Open source home automation that puts local control and privacy first.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
PyTorch native quantization and sparsity for training and inference
Neptune Fetcher is designed to separate data retrieval capabilities from the regular neptune package. This separation makes data fetching more efficient and improves performance.
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
A set of experiments using PyLaia on different datasets
Logging module with nicer formatting
A native PyTorch Library for large model training
Pretrain, finetune, deploy 20+ LLMs on your own data. Uses state-of-the-art techniques: flash attention, FSDP, 4-bit, LoRA, and more.
Essential guides and programming tools in my toolbox (with focus on ML Training)
Blazingly fast, distributed streaming of training data from any cloud storage for training AI models
Source to source compiler for PyTorch. It makes PyTorch programs faster on single accelerators and distributed.
A framework for few-shot evaluation of autoregressive language models.
A Fusion Code Generator for NVIDIA GPUs (commonly known as "nvFuser")
NeurIPS Large Language Model Efficiency Challenge: 1 LLM + 1GPU + 1Day
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper GPUs, to provide better performance with lower memory utilization in both training and inference.
Outlining techniques for improving the training performance of your PyTorch model without compromising its accuracy
Enabling PyTorch on Google TPU
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
UVA programming challenges
Build and train PyTorch models and connect them to the ML lifecycle using Lightning App templates, without handling DIY infrastructure, cost management, scaling, and other headaches.
A latent text-to-image diffusion model
Taming Transformers for High-Resolution Image Synthesis
Ongoing research training transformer models at scale
FFCV: Fast Forward Computer Vision (and other ML workloads!)
My java solutions of the programming puzzles.
CPU & CUDA implementation of several neural network utils