VibhuJawa/NeMo-Curator
Scalable toolkit for data curation
Senior Software Engineer @Nvdia | Former CS Grad student at Johns Hopkins University
Scalable toolkit for data curation
WebMainBench is a high-precision benchmark for evaluating web main content extraction.
cuDF - GPU DataFrame Library
A high-throughput and memory-efficient inference and serving engine for LLMs
cuML - RAPIDS Machine Learning Library
ArcticInference: vLLM plugin for high-throughput, low-latency inference
Integration between Lance and Ray for distributed data processing
Rapids Analytics Framework Toolset to share building blocks between cuGraph and cuML
This repo implements attention networks for visual question answering
NVIDIA Ingest is an early access set of microservices for parsing hundreds of thousands of complex, messy unstructured PDFs and other enterprise documents into metadata and text to embed into retrieval systems.
A fast programming language detector
OWASP Top 10 for Large Language Model Apps (Part of the GenAI Security Project)
Metric calculation library
A lightweight data processing framework built on DuckDB and 3FS.
Utilities for Dask and CUDA interactions
This repo benchmarks pytorch vs tensorflow for speech recognition
cuGraph - RAPIDS Graph Analytics Library
This repository demonstrates how to efficiently create embeddings at unparalleled speeds using RAPIDS, PyTorch, and Sentence Transformers. The project aims to maximize GPU efficiency to accelerate the process of embedding generation, with robust support for multi-node multi-GPU (MNMG) setups
PDF GPT allows you to chat with the contents of your PDF file by using GPT capabilities. The only open source solution to turn your pdf files in a chatbot!
This repo contains the DGL cugraph Examples
Python package built to ease deep learning on graph, on top of existing DL frameworks.
RAPIDS GPU-BDB
A distributed task scheduler for Dask
Build environments for various dask related projects on gpuCI