mfuntowicz/lmir
Language Model Intermediate Representation
@huggingface | Low-level stuff
Language Model Intermediate Representation
HMLL - High-Performance Model Loading Library for Efficient AI Model I/O
LLM inference in C/C++
Experiment tracking for hmll repository - exploring not so clear paths and ideas that leads certainly nowhere
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
The fastest way to install llama.cpp
Lightning Fast LLM Tokenizer
Accessible large language models via k-bit quantization for PyTorch.
Pure C based implementation of safetensors deserialization
"AI-Trader: Can AI Beat the Market?" Live Trading: https://hkuds.github.io/AI-Trader/
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Turn your Raspberry Pi into a powerful weather station
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
PyTorch RNet implementation with Distributed and Mixed-Precision training support.
2D Cutting Problem solution generator using metaheuristic methods and linear programming
Automatically generates Rust FFI bindings to C (and some C++) libraries.
AWS Deep Learning Containers (DLCs) are a set of Docker images for training and serving models in TensorFlow, TensorFlow 2, PyTorch, and MXNet.
A PyTorch Extension: Tools for easy mixed precision and distributed training in Pytorch
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
The Triton backend for the ONNX Runtime.
Open standard for machine learning interoperability
Flax is a neural network library for JAX that is designed for flexibility.
Fast BPE