nelson-liu/lost-in-the-middle
Code and data for "Lost in the Middle: How Language Models Use Long Contexts"
@stanfordnlp CS PhD student.
Code and data for "Lost in the Middle: How Language Models Use Long Contexts"
A quick intro to Cython for Python users who don't know C
Scripts for fine-tuning Llama2 with composable FSDP & PEFT methods to cover single/multi-node GPUs. Supports default & custom datasets for applications such as summarization & question answering. Supporting a number of candid inference solutions such as HF TGI, VLLM for local or cloud deployment.Demo apps to showcase Llama2 for WhatsApp & Messenger
Code for the paper "Discovering Phonesthemes with Sparse Regularization", to be presented at the NAACL 2018 Workshop on Subword and Character Level Models in NLP.
Various models and code (Manhattan LSTM, Siamese LSTM + Matching Layer, BiMPM) for the paraphrase identification task, specifically with the Quora Question Pairs dataset.
Companion repo for "Evaluating Verifiability in Generative Search Engines".
Code for the paper "Inoculation by Fine-Tuning: A Method for Analyzing Challenge Datasets", to be presented at NAACL 2019.
A toolkit for evaluating the linguistic knowledge and transferability of contextual representations. Code for "Linguistic Knowledge and Transferability of Contextual Representations" (NAACL 2019).
A fast, simple and lightweight Bloom filter library for Python, implemented in Rust.
Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.
Streaming WARC/ARC library for fast web archive IO
A crate for uploading files to Google cloud storage, and for generating download urls.
Dump the text of the Gigaword dataset into a single file, for use with language modeling (and other!) toolkits
A collaborative platform for reproducible research (web interface and CLI).
🏎 A set of primitives to build simple, flexible, WAI-ARIA compliant React autocomplete, combobox or select dropdown components.
🤗 Transformers: State-of-the-art Natural Language Processing for TensorFlow 2.0 and PyTorch.
A simple HTML content extractor in Python. Can be run as a wrapper for Mozilla's Readability.js package or in pure-python mode.
💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
A robust web archive analytics toolkit
Given a description of a color, return its closest standard HTML4 color.
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
Helps you write algorithms in PyTorch that adapt to the available (CUDA) memory
A library to analyze PyTorch traces.
A high-throughput and memory-efficient inference and serving engine for LLMs
Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more
A set of Python scripts for preprocessing the Wikidata JSON dump and running simple queries in an efficient manner.
Painless hash link routing for React applications.