Eylon Eliyahu Krause

@EylonKrause · User

GitHub profile ↗ · Compare

Sr. HW/FW Algorithm Engineer. This page contains code from/for: GPU algorithms and communications, published papers, projects and present/former lecturing.

Tel Aviv - Israel 2 followers49 repositories

Repositories

EylonKrause/Windows-11-in-Assembly

Windows 11 in Assembly — aggressive, binary-level superoptimization of a live Windows 11 install, validated change by change

★ 1CForks 0

EylonKrause/Simulating-ICIs-in-embedded-C

200 Gb/s/lane PAM4 SerDes lane in C17: channel, AFE, equalisers, CDR, and the fixed-point firmware that controls them through a register HAL. Unit-tested with no hardware.

★ 0CForks 0

EylonKrause/ml_dtypes

A stand-alone implementation of several NumPy dtype extensions used in machine learning.

★ 0Forks 0

EylonKrause/STL

MSVC's implementation of the C++ Standard Library.

★ 0Forks 0

EylonKrause/shiftfly

A shift-routed interconnect for accelerator fabrics beyond pod scale. Audits every Google TPU topology against the Moore bound; replaces Boardfly's fully-connected global tier with a generalized Kautz digraph on the OCS. Two-sided results, stdlib only.

★ 0PythonForks 0

EylonKrause/robust-accelerator-tradeoffs

Where does the next watt go? Choosing an AI-accelerator budget allocation when the workload mix is unknown -- max-expected vs minimax-regret over the whole forecast space. Normalized model, stdlib only.

★ 0PythonForks 0

EylonKrause/Surelog

SystemVerilog 2017 Pre-processor, Parser, Elaborator, UHDM Compiler. Provides IEEE Design/TB C/C++ VPI and Python AST & UHDM APIs. Compiles on Linux gcc, Windows msys2-gcc & msvc, OsX

★ 0Forks 0

EylonKrause/verible

Verible is a suite of SystemVerilog developer tools, including a parser, style-linter, formatter and language server

★ 0Forks 0

EylonKrause/iree

A retargetable MLIR-based machine learning compiler and runtime toolkit.

★ 0Forks 0

EylonKrause/xla

A machine learning compiler for GPUs, CPUs, and ML accelerators

★ 0Forks 0

EylonKrause/XNNPACK

High-efficiency floating-point neural network inference operators for mobile, server, and Web

★ 0Forks 0

EylonKrause/highway

Performance-portable, length-agnostic SIMD with runtime dispatch

★ 0Forks 0

EylonKrause/cuPythia

CUDA Pythia: GPU-acceleration experiments + audit on Pythia 8.317

★ 0C++Forks 0