KshitijLakhani/VLSI_Verilog_Projects
RTL Synthesis for Fast Arithmetic circuits like Booth encoded Multipliers, Carry Save Adders, Fixed-Point and Floating-Point conversions, Control circuits like State Machines, and DSP applications like FFT.
Deep Learning Performance Engineer @ NVIDIA. MS in ECE @ UC Davis. Interested in Machine Learning and Parallel Computing.
RTL Synthesis for Fast Arithmetic circuits like Booth encoded Multipliers, Carry Save Adders, Fixed-Point and Floating-Point conversions, Control circuits like State Machines, and DSP applications like FFT.
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper and Ada GPUs, to provide better performance with lower memory utilization in both training and inference.
A simple, performant and scalable Jax LLM!
Matrix multiplication is used to investigate and optimize blocking factor, associativity and size of cache hierarchy. Attributes like miss rate and execution time are tabulated and compared.
Efficient Implementation of CUDA Algorithms like scan, reduce, histogram, sort and search among others via applications like blurring of images, red eye removal, searching and histogram equalization
Implementation of some basic Machine Learning Tasks like logistic and linear regression, simple image classification, Deep CNNs, etc. using python and keras primarily
Deep CNN model using data augmentation with a very high accuracy of 78.51% on Kaggle's fer2013 dataset
Trial space