holtwood/LessUp

Personal Project Hub: AI Infrastructure & High-Performance Computing Engineering Practice | 个人项目主页:AI 基础设施与高性能计算工程实践

★ 3Forks 0GitHub ↗Compare
ai-infrastructurecudagpu-computinghpcprofile

README

Static title

聚焦 AI 基础设施、CUDA Kernel 与高性能系统工程

🔬 Focus: AI Infrastructure · CUDA Kernels · LLM Inference · HPC Systems
🌱 Currently: Building high-throughput inference pipelines and GPU-first systems
🤝 Open to: AI infrastructure, performance engineering, research collaboration, and open-source collaboration


Followers   Stars   Views



Profile  Selected Work  Background  Stack  Signals  Connect



👨‍💻 About Me / 关于我

Top Languages

I build AI infrastructure and GPU-first high-performance systems with C++/CUDA, Python, and Go. 主要聚焦 AI 基础设施、GPU 算子优化与高性能系统工程实践。

  • 🔥 GPU Kernel Engineering — CUDA/Triton kernels for FlashAttention, GEMM, quantization, and memory-aware operator design
    GPU 算子工程 — FlashAttention、GEMM、量化与内存感知算子设计
  • 🧠 AI Inference Systems — lightweight LLM runtimes, KV Cache, W8A16/FP8 quantization, and inference path optimization
    AI 推理系统 — 轻量 LLM 运行时、KV Cache、量化方案与推理路径优化
  • ⚡ High-Performance Computing — simulation, rendering, and image-processing pipelines tuned for throughput and scalability
    高性能计算 — 面向吞吐与可扩展性的仿真、渲染与图像处理流水线
  • 🌐 Real-time Systems — RTC signaling, streaming applications, and digital human platforms with system-level integration
    实时系统 — RTC 信令、流媒体应用与数字人平台的系统级集成

Currently / 当前关注: inference acceleration, kernel fusion, and end-to-end GPU system design.
推理加速、算子融合与端到端 GPU 系统设计。



🚀 Selected Work / 项目全景

Featured Projects / 核心项目 — Start here for the quickest overview of my work in bioinformatics, HPC, AI inference, and developer tooling.
如果你想快速判断我的技术重心与代表作,建议先看下面 4 个项目。
Best entry points for collaboration, hiring conversations, and technical review.

High-performance FASTQ compression with 3.97x ratio and O(1) random access. C++23, ABC+SCM algorithms.
高性能 FASTQ 压缩:3.97x 压缩比,O(1) 随机访问,C++23 + oneTBB。

Stars C++23

High-performance FASTQ QC toolkit (stat/filter/trim); zero-copy I/O, TBB pipeline, C++23.
高性能 FASTQ 质控工具:零拷贝 I/O、TBB 流水线、C++23。

Stars C++23

End-to-end Metagenomic Intelligence and Comprehensive Omics Suite (Mammoth Cup 2024)
端到端宏基因组综合分析平台(猛犸杯 2024 参赛项目)

Stars R

Systematic knowledge base for bioinformatics (Chinese community)
面向中文社区的生物信息学体系化知识库

Stars MDX

🚀 AI Infra Portfolio / AI Infra 五仓作品集(open-infra-ai)

Original, sole-authored engineering portfolio — from CUDA kernels to a working inference engine, with tests, benchmarks & honest measurement records.
从 CUDA 算子到可运行推理引擎的个人原创作品集(含测试、基准与诚实记录的测量口径)。入口:open-infra-ai/aicl-lab

CUDA operator engineering path: SGEMM ladder to reusable inference components
CUDA 算子工程学习路径(SGEMM 优化阶梯)

FlashAttention fwd/bwd from scratch in CUDA C++ (FP16/BF16 WMMA)
从零实现的 CUDA FlashAttention 前后向

Triton kernel library (RMSNorm+RoPE / SwiGLU / FlashAttn / SGEMM) + torch.library registration
Triton 算子库 + torch.library 注册

CUDA-native C++ inference engine: GGUF loading, W8A16 quantization, paged-KV strategy; 170 tests
C++ 原生推理引擎:GGUF / W8A16 / 分页 KV

PagedAttention-style paged KV + continuous batching control plane (Rust), e2e-verified vs llama.cpp
分页 KV + Continuous Batching 控制面(Rust)

Portfolio landing page: evidence pack, benchmarks methodology & interview storytelling
作品集入口:证据包、基准口径与项目讲述

🎯 AI Infra Career Transition / AI Infra 转行计划

12-week AI Infra transition plan (2026-08-24 ~ 2026-11-15): skill matrix, weekly plans, interview matrix & application pipeline.
12 周 AI Infra 转行计划:能力矩阵、逐周计划、面试矩阵与求职执行

CUDA LLM Inference

Standalone navigation center: full repo inventory, fork/upstream audit, migration log & deep-dive guides.
独立仓库导航中心:全量盘点、Fork 上游审计、迁移记录与深读指南

AI Infra Repository Index

🧬 Bioinformatics & Genomics / 生物信息学

Systematic knowledge base for bioinformatics (Chinese community)
面向中文社区的生物信息学体系化知识库

MDX Bioinformatics

End-to-end Metagenomic Intelligence and Comprehensive Omics Suite
端到端宏基因组综合分析平台(猛犸杯 2024)

R Metagenomics

High-performance FASTQ compression with 3.97x ratio and O(1) random access. C++23, ABC+SCM.
高性能 FASTQ 压缩:3.97x 压缩比,O(1) 随机访问

C++23 oneTBB

High-performance FASTQ QC toolkit (stat/filter/trim); zero-copy I/O, TBB pipeline, C++23.
高性能 FASTQ 质控工具:零拷贝 I/O、TBB 流水线、C++23

C++23 Zero-Copy

Curated bioinformatics algorithms knowledge base with complexity analysis, CLI tools, and bilingual docs.
精选生物信息学算法知识库,含复杂度分析、CLI 维护工具与双语文档

Python Algorithms

⚡ CUDA & HPC / 高性能计算

CUDA image processing experiments: parallel filters and pipelines on GPU.
CUDA 图像处理实验:GPU 并行滤波与流水线

CUDA Image Processing

High-performance C++ optimization guide with lock-free data structures, SIMD, and memory optimization.
高性能 C++ 优化指南,含无锁数据结构、SIMD 和内存优化示例

C++17 SIMD

Header-only C++23 bit manipulation library with SIMD acceleration (SSE2/AVX2/AVX-512/NEON).
仅头文件 C++23 位操作库,支持 SIMD 加速

C++23 SIMD

Classic lossless compression algorithms in C++17, Go, and Rust with cross-language binary verification.
经典无损压缩算法,支持 C++17、Go 和 Rust,跨语言二进制验证

Go Rust

My solutions to TensorTonic problems (tensor/GPU computing drills).
TensorTonic 题解(张量与 GPU 计算练习)

Python Tensor

Compression Knowledge Base: Algorithm Theory, Performance Benchmarks & C++ Examples
压缩算法知识库:原理、性能基准与 C++ 示例

C++17 Algorithms

🤖 AI & Developer Tooling / AI 与开发者工具

Cursor AI 编程规则精选集 | 132+ 规则,覆盖前端/后端/AI/DevOps 等 32 个领域

Stars JavaScript

Local-first bookmark manager with structured storage and import/export.
本地优先的书签管理器:结构化存储与导入导出

TypeScript Local-first

Notes syncing utility built around Brave browser data.
围绕 Brave 浏览器数据构建的笔记同步工具

JavaScript Tooling

Offline-first bookmark cleaner: rules-first, ML-assisted, LLM-optional
智能书签清理与分类:规则+ML+LLM(可选)

Python ML

Multi-Model Real-Time Visual Recognition System with REST API and WebSocket Streaming
多模型实时视觉识别系统,提供 REST API 和 WebSocket 流式推理

Python YOLOv8

Privacy-first diagram editor with local WASM rendering, Kroki full mode, sharing, and export.
隐私优先的图表编辑器:本地 WASM 渲染、Kroki 全模式、分享与导出

TypeScript WASM

🌐 Applications / 应用项目

Browser-native 3D digital human engine with voice, vision & dialogue. Zero-config, offline-ready.
浏览器原生 3D 数字人引擎,支持语音、视觉与对话。零配置、离线可用。

Stars TypeScript

Lightweight WebRTC Demo: Go Signaling Server + Vanilla JavaScript Client, OpenSpec-Driven
轻量级 WebRTC 演示:Go 信令服务 + 原生 JavaScript 客户端,OpenSpec 驱动开发

Go WebRTC

Browser-based memory training PWA with FSRS-4.5 spaced repetition, N-back training, and adaptive difficulty
基于 FSRS-4.5 间隔重复、N-back 训练和自适应难度的浏览器记忆力训练 PWA

JavaScript PWA


🎓 Background & Experience / 教育与经历

🎓 Education

Xidian University Xidian University

Computer Science related background. / 计算机科学相关背景

💼 Experience

Mindray Mindray · ZEGO ZEGO · BGI BGI

Engineering across medical imaging, RTC systems, and genomic-scale data workflows. / 覆盖医疗影像、实时音视频系统与基因数据工程。


🛠️ Tech Stack / 技术栈

聚焦与核心项目强相关的技术:AI Infrastructure · CUDA Kernel Engineering · LLM Inference · HPC Systems

Category Technologies
AI Infrastructure AI Infrastructure
CUDA Kernel Engineering CUDA C++ Tensor Core WMMA FlashAttention
LLM Inference Optimization Triton TensorRT KV Cache W8A16 / FP8
HPC Performance Engineering SIMD Intel oneTBB Memory Optimization Performance Profiling

📊 Signals & Activity / 数据概览

LessUp's GitHub stats

GitHub Activity Graph (public, UTC)
Public data only · UTC aggregation · may lag a few hours / 仅统计公开贡献、按 UTC 聚合,可能有数小时延迟

📫 Collaboration & Contact / 联系方式

Reach out if you're building AI infrastructure, inference acceleration, GPU systems, or performance-critical tooling.
欢迎联系我交流 AI 基础设施、推理加速、GPU 系统,以及对性能敏感的工程项目。
Open to technical collaboration, engineering roles, research discussions, and thoughtful open-source work.
Email   GitHub

Footer

Contributors

holtwood

Issues