Jokeren/Notes

Computer Science Reading Notes

★ 8Forks 4GitHub ↗Compare

README

Notes

Have fun reading papers and books!

Concurrency

  1. Read-Log-Update A Lightweight Synchronization Mechanism for Concurrent Programming
  2. A Pragmatic Implementation of Non-Blocking Linked-Lists
  3. A Wait-free Queue as Fast as Fetch-and-Add
  4. Non-blocking Patricia Tries with Replace Operations
  5. The Java Memory Model
  6. Foundations of the C++ Concurrency Memory Model
  7. The Foundations for Scalable Multi-core Software in Intel® Threading Building Blocks
  8. SALSA: Scalable and Low Synchronization NUMA-aware Algorithm for Producer-Consumer Pools
  9. NUMASK: High Performance Scalable Skip List for NUMA
  10. BRAVO–Biased Locking for Reader-WriterLocks
  11. Eraser: A Dynamic Data Race Detector for Multithreaded Programs
  12. ThreadSanitizer – data race detection in practice
  13. Scheduling Multithreaded Computations by Work Stealing
  14. Wait-Free Synchronization
  15. CDSCHECKER Checking Concurrent Data Structures Written with CC++ Atomics
  16. Algorithms for Scalable Synchronization on SharedMemory Multiprocessors
  17. Everything You Always Wanted to Know About Synchronization but Were Afraid to Ask
  18. Nonblocking Concurrent Data Structures with Condition Synchronization
  19. Software Transactional Memory for Dynamic-sized Data Structures
  20. Transactional Memory
  21. Transactional Data Structure Libraries

Parallel Computing

  1. A programming system for future proofing performance critical libraries
  2. BLIS A Framework for Rapidly Instantiating BLAS Functionality
  3. Barrier Elision for Production Parallel Programs
  4. Implementing Strassen's Algorithm with BLIS
  5. Lightweight Dynamic Selection for Kernel-based Data-parallel Programming Model
  6. The Big Data Challenges of Connectomics
  7. Experimenting with Low-Overhead OpenMP Runtime on BG/Q
  8. Reducers and Other Cilk++ Hyperobjects

Deep Learning

  1. Deep Learning with Limited Numerical Precision
  2. One weird trick for parallelizing convolutional neural networks
  3. Overcoming Challenges in Fixed Point Training of Deep Convolutional Networks
  4. Scaling Distributed Machine Learning with the Parameter Server
  5. Optimizing Memory Efficiency for Deep Convolutional Neural Networks on GPUs

GPUs

  1. Efficient Synchronization Primitives for GPUs
  2. Enterprise Breadth-First Graph Traversal on GPUs
  3. GPU Multisplit
  4. iBFS Concurrent Breadth-First Search on GPUs.
  5. Analyzing CUDA Workloads Using a Detailed GPU Simulator
  6. Demystifying GPU Microarchitecture through Microbenchmarking
  7. NVIDIA TESLA V100 GPU ARCHITECTURE
  8. Visualizing Complex Dynamics in Many-Core Accelerator Architectures

CPUs

  1. Simultaneous Multithreading Maximizing On-Chip Parallelism
  2. A Single-Chip Multiprocessor
  3. Evaluating the Potential of Multithreaded Platforms for Irregular Scientific Computations
  4. Victim Replication: Maximizing Capacity while Hiding Wire Delay in Tiled Chip Multiprocessors
  5. The future of microprocessors
  6. Elastic cooperative caching: an autonomous dynamically adaptive memory hierarchy for chip multiprocessors
  7. IBM POWER7 multicore server processor
  8. A Primer on Memory Consistency and Cache Coherence
  9. IBM Blue Gene/Q memory subsystem with speculative execution and transactional memory
  10. Speculative Lock Elision: Enabling Highly Concurrent Multithreaded Execution

Machine Learning

  1. A Few Useful Things to Know about Machine Learning

Compiler

  1. Graspan: A Single-machine Disk-based Graph System for Interprocedural Static Analyses of Large-scale Systems Code
  2. Program Locality Analysis Using Reuse Distance

Contributors

Jokeren

Issues