ayerofieiev-tt/cloud-native-stack-intro
Short explainer film: from single host to cloud-native (Remotion)
Software @tenstorrent
Short explainer film: from single host to cloud-native (Remotion)
Single-concurrency latency benchmark for OpenAI-compatible /v1/chat/completions endpoints
Interactive zoomable tour of the Tenstorrent Blackhole chip — from the 17×12 tile grid down to a single SFPU lane.
SGLang is a high-performance serving framework for large language models and multimodal models.
Post-training inference optimization: quantization, pruning, distillation, speculative decoding
TT-Metal take-home: single-core tile addition on ttsim simulator
Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs
The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.
A high-performance C++ machine learning framework with lazy evaluation, designed for fast dispatch times and efficient computation graphs.