dublinsubway/GHAScanCombinator

★ 0Forks 0PythonGitHub ↗Compare

README

GHAScanCombinator

A tool that combines four (yes, four!) different GitHub Actions' workflow analyzers and puts results into GitHub issues (and a CSV file, if it's more to your liking).

Tools used

zizmor, octoscan, semgrep and github-actions-scanner.

zizmor

Installed via pip.
Command used: zizmor --format=sarif --no-exit-codes --offline -q --pedantic [file_path]

octoscan

Warning

Octoscan outputs an error code when it gets a finding, so error code is ignored there when it's being run. Check logs.

Complied with go build from a repo.
Command used: ../tools/octoscan/octoscan scan [file_path] --format sarif -v

semgrep

Installed via pip.
Command used: semgrep scan [directory_path] --no-autofix --no-error -q --sarif-output=[output_path]

github-actions-scanner

Installed via cloning the repo and running npm install.
Command used: npm run start -- scan-org -o [organization_name] --format json --output [output_path]

Layout

scripts: Start here. Code used to run the entire process is located there.
tools: Where tools are meant to be installed, which are used for the scanner.
data: Where data gathered from repository crawling is stored. Separated by folders relating to each project.
results: Results after tools have been executed. Folder csv contains the summarised results.

Setup steps

Prerequisites

python3, pip
Ability to create python virtual environments (venv)
Latest node (22.15.0 worked for me) and npm (0.40).
Go compiler
Shell
GitHub token
GitHub CLI

The setup process itself

  1. Clone this.

  2. Add your GitHub token to .env file, as shown in .env.example; if you want to use not the dotcom (www.github.com) instance of GitHub, provide the instance URL as well.

  3. Go to scripts folder in the terminal in your console (needed so paths are resolved properly)

  4. Create a file with organization names that you want to analyze (new line each).
    Example (also given in sample_org_names.txt):

     org1
     org2
     org3
    
  5. Set install.sh to be executable and run it (changes files perms to be executable as well; if throws errors, try running as sudo). When over, start the analysis by launching run.sh <name_of_file_with_orgs>.

  6. Make yourself some tea and coffee.

  7. Find a list of results in results/csv folder.

  8. If you desire to create issues in your project with all the findings, login to GitHub CLI (gh auth login) and run:

     source venv/bin/activate
     python3 create_issues.py
    
  9. Alternatively, to show some predefined charts based on your data, run:

     source venv/bin/activate
     python3 visualize_data.py
    

That's it!

Caveats

  • Limit is 60 requests per hour without GH token and 5000 with token.
  • Tested and created on Ubuntu 22.04; may have issues when used on Windows.
  • There no pagination in get-all-repos function (which gets repos from org for cloning), as it wasn't needed in this case; only supports 100 repos or less for now.
  • The way script decides whether the file is a workflow or not is by checking file extension; some of the cloned files might be, eg, kubernetes files instead.

Licensing

This project includes code/components which is licensed under Apache 2.0 (github-actions-scanner), GNU Lesser General Public License v2.1 (semgrep), GNU General Public License v3.0 (octoscan) and MIT (zizmor).

This project itself is licensed under MIT license.

Contacts

Open a GitHub issue here.

Contributors

dublinsubway

Issues