BartleyR/ubuntu-setup

Instructions on how to configure a Ubuntu install like I do for GPU DS/ML/DL

★ 0Forks 0GitHub ↗Compare

README

Ubuntu 20.04 Workstation Setup

Introduction

I use my GPU-enabled workstation to run data science, ML, and DL workflows using a large number of open-source packages/projects. This includes, but is not limited to, RAPIDS, PyTorch, TensorRT, TensorFlow, multiple visualization packages, and much of the PyData ecosystem. I primarily develop using Docker containers, so the minimalist approach to setup here reflects that. These instructions are manual. If you prefer something more automated, my colleague Paul has a repo that automates setup of a new Xubuntu install.

Setup Instructions

  1. Install Ubuntu 20.04

  2. Restart (boot into recovery mode if you’re having display issues)

  3. Install GCC/G++

    sudo apt install gcc
    
  4. Install NVIDIA driver and CUDA toolkit

    wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/cuda-ubuntu2004.pin
    sudo mv cuda-ubuntu2004.pin /etc/apt/preferences.d/cuda-repository-pin-600
    sudo apt-key adv --fetch-keys https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/7fa2af80.pub
    sudo add-apt-repository "deb https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/ /"
    sudo apt-get update
    sudo apt-get -y install cuda
    
  5. Install and configure SSH

    1. Install SSH

      sudo apt install openssh-server
      
    2. Configure your client to connect to the workstation via SSH

      1. Instructions on how to setup SSH keys and connect with keys [recommended]
      2. Insturctions on how to connect via IP and password
  6. Restart [optional, but gets you out of recovery mode]

  7. Mount additional drives [optional]

    1. Edit /etc/fstab (detailed instructions on how to find UUID and edit)
    2. Run sudo mount -a to reflect edited changes
  8. Relink all folders with symbolic links [optional]

    # For every folder you want to symlink
    ln -s <path/to/source/folder> <path/to/target>
    
  9. Install git CLI tools

    sudo apt install git
    
  10. Install curl

    sudo apt install curl
    
  11. Configure ZSH and set as default shell [optional, skip if you want to keep Bash]

    1. Install ZSH

      sudo apt install zsh
      
    2. Install custom dotfiles from GitHub repo

    3. Change default shell to ZSH

      chsh -s $(which zsh)
      
    4. Modify path to include link to CUDA binaries

      # Add to ~/.aliases_functions.local
      path=('/usr/local/cuda/bin' $path)
      export PATH
      
  12. Install Docker CE

    # Install dependencies
    sudo apt install \
        apt-transport-https \
        ca-certificates \
        curl \
        gnupg-agent \
        software-properties-common
    
    # Add Docker GPG key
    curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo apt-key add -
    
    # Setup Docker stable repo
    sudo add-apt-repository \
       "deb [arch=amd64] https://download.docker.com/linux/ubuntu \
       $(lsb_release -cs) \
       stable"
    
    # Install Docker engine
    sudo apt update
    sudo apt install docker-ce docker-ce-cli containerd.io
    
  13. Setup Docker to manage it as a non-root user

    # Add the docker group (might already exist)
    sudo groupadd docker
    
    # Add your account to the docker group
    sudo usermod -aG docker $USER
    
    # Need to log out/in to reevaluate group membership
    
  14. Setup NVIDIA container toolkit

    # Setup stable repo and GPG key
    distribution=$(. /etc/os-release;echo $ID$VERSION_ID) \
       && curl -s -L https://nvidia.github.io/nvidia-docker/gpgkey | sudo apt-key add - \
       && curl -s -L https://nvidia.github.io/nvidia-docker/$distribution/nvidia-docker.list | sudo tee /etc/apt/sources.list.d/nvidia-docker.list
    
    # Install nvidia-docker2 and dependencies
    sudo apt update
    sudo apt install -y nvidia-docker2
    
    # Restart Docker daemon
    sudo systemctl restart docker
    
    # Test by running a base CUDA container
    docker run --rm --gpus all nvidia/cuda:11.0-base nvidia-smi
    
  15. Verify NVLink is functioning [optional]

    # Copy CUDA samples to home directory
    cuda-install-samples-11.1.sh ~/
    
    # Get in the right directory
    cd ~/NVIDIA_CUDA-11.1_Samples/0_Simple/simpleP2P/
    
    # Compile the sample
    make
    
    # Run the test
    ./simpleP2P
    
    # Can also run the bandwidth latency test
    cd ~/NVIDIA_CUDA-11.1_Samples/1_Utilities/p2pBandwidthLatencyTest
    make
    ./p2pBandwidthLatencyTest
    
  16. Install AWS CLI tools

    # Download AWS CLI file
    curl "https://awscli.amazonaws.com/awscli-exe-linux-x86_64.zip" -o "awscliv2.zip"
    
    # Extract file
    unzip awscliv2.zip
    
    # Install AWS CLI
    sudo ./aws/install
    

Optional Next Steps

Sometimes (rarely these days) I might need multiple CUDA toolkit versions on the bare metal OS. For this latest install, I'm skipping this and going to try getting by with CUDA 11.x on the host OS while managing other CUDA toolkits via containers (if necessary). If you need multiple CUDA toolkit versions on your install, Paul's script on how to configure this is very useful.

Acknowledgments

I relied heavily on Paul's Xubuntu bootstrap scripts to help make this simple walkthrough.

Contributors

BartleyR

Issues