Hardening a GitHub Actions Workflow
When it comes to my personal website I want something with as few moving parts as possible, and that made Hugo a simple choice to make when I first built this site several years ago. Hugo uses Markdown to generate static HTML that can be served by good old-fashioned nginx without needing a database or anything.
Historically I would author content locally, build the site with Hugo, and then SFTP everything up to the server to deploy it, but with some spare time on my hands recently I decided to move this to a GitHub Actions (GHA) deploy-on-commit workflow. And it works just great! But as you surely already know, more moving parts means more attack surfaces - so let’s think through how we might secure such a thing using Adam Shostack’s four questions.
The Threat Model#
What are we doing? We’re automating how my personal site gets deployed. It used to be the case that I would deploy by just SFTP’ing the Hugo-built HTML files to the server and call it a day: now, when a new commit lands in main for the site’s repo a GHA workflow handles the build with Hugo and deploys it to the server. Great!
But what could go wrong? Quite a few things:
- Supply chain vulnerabilities. Everyone’s doing it! Adding a remote build step to the mix means that there is more of a reliance on third-party packages (GitHub’s own packages for GHA, Hugo running in a remote environment) and all of those packages are themselves attack surfaces in the modern era.
- Access to the server itself. The entire access model up until now was SSH keys on my personal laptop, which largely limits the attack surface to something like the wrench model. To deploy automatically, I need to give someone (or something) access to my actual production server. And that implies the third issue:
- Secrets management. We need to think about how we will be securely storing (and sharing) the various secrets - namely, SSH keys - that will allow trusted processes to connect to my running servers.
- Compromised Hugo builds produce output that I do not want on my site.
- Someone compromises my accounts.
If an attacker were to compromise any part of this the principal risk is that they could put whatever content they want on my website, such as crypto scams or harmful messages that suggest I only deadlift with straps (NOT TRUE). That’s not really a huge deal in the grand scheme of things as my personal blog does not for example direct air traffic or handle payments - but it’s obviously something we want to prevent as security-conscious folks.
So, what are we going to do about all that?
Supply Chain Vulnerabilities#
The vector that I’m actually most concerned with is that a package I rely on gets compromised and some malware gets pulled in - this is a thing that happens basically all the time now, and probably should keep you awake at night if your company is dependent on 9,000 little NPM packages of dubious provenance. So a good first step is to limit how many packages we’re dependent on in the first place.
To that end, we really only need a few things:
- Hugo itself (to build the site). This is an example of the security/convenience tradeoff: I could remove this dependency entirely from the build environment and continue to build my site with the locally-installed version. But it’s really nice to automate that piece.
- Using GitHub Actions for this means we depend on some of the packages that GitHub publishes under the
actionsproject:checkoutto actually check out the website’s code and do things with it, and thenupload-artifactanddownload-artifactas a means of storing and retrieving the built site. More on this momentarily.
GitHub runs a very good security team and so I am not super concerned with actions/* being compromised. Regardless, we can balance staying up-to-date with regard to security fixes and avoiding getting caught out by a compromised install by setting a cooldown so that a package needs to have been published for some amount of time before Dependabot suggests an update, and pinning to specific SHAs so that all of our builds use the exact versions we request. Hugo is not managed by Dependabot and is downloaded to the build environment each deploy, and so we instead store a checksum of our current version that we check on execution (if the checksum doesn’t match what we just downloaded, then we bail out) and that only gets updated whenever I think to do it.
You can define a Dependabot cooldown in .github/dependabot.yml:
- package-ecosystem: github-actions
directory: /
schedule:
interval: weekly
cooldown:
default-days: 7
And you can stitch together a checksum check for a Hugo binary like so:
env:
HUGO_VERSION: 0.166.0
HUGO_SHA256: <the checksum of that version>
steps:
- name: Install Hugo (checksum-verified, no root)
run: |
archive="$RUNNER_TEMP/hugo.tar.gz"
curl -fsSL -o "$archive" \
"https://github.com/gohugoio/hugo/releases/download/v${HUGO_VERSION}/hugo_extended_${HUGO_VERSION}_linux-amd64.tar.gz"
echo "${HUGO_SHA256} ${archive}" | sha256sum --check --strict -
Finally, pinning an action version to a SHA looks like this:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
One note about pinned SHAs for dependencies is the need to verify that they actually belong to the source you’re trying to get them from. GitHub by design will check forks of a repo when attempting to resolve a particular SHA, and maintainers can be caught out by a PR that mentions “updated dependencies” and is pointing to a fork that contains malicious code. This is not really a concern for me because there aren’t going to be any PRs made to this repo, but if you’re doing this for Real Code, it’s worth keeping in mind (zizmor, a static analyzer for GHA workflows, has an impostor-commit audit rule that can detect this specific problem).
Secrets#
There are a surprising number of quirks around secrets management in GitHub to be aware of. One of the biggest is that repository-level secrets cannot be scoped. Any workflow, in any job, on any branch can ask for any of them, and once a job has a secret, every step in that job can get at it. The compromise of tj-actions/changed-files is an example of this: the malicious code just dumped the GHA runner’s memory into the (public for public repos!) workflow logs.
This is a problem for our setup: if we assume that some breach eventually happens in one of our dependencies then we want to contain it as much as possible and that means a clean separation between the “attackers can get a foothold by compromising a package” environment and the “has secrets that could be used to access the prod server” environment. Luckily, there’s a nicely simple fix for this:
- Install Hugo and build the website in its own job and then upload the built HTML/CSS.
- Download the artifact (the built website) in another job that deploys it to prod, and has access to the secrets it needs.
- Use environment-level secrets rather than repository-level. Store the actual secrets in a
productionenvironment and only declare that environment (in the workflow YAML, that is theenvironment: productiondirective) in the deploy job, which needs the secrets. The Hugo build job declares no environment at all and runs on a separate machine so it never sees those secrets. The production environment additionally only accepts runs frommain.
It’s technically the case that anyone with commit access could define other jobs that run in the production environment, but “anyone with commit access” includes:
- Me
- Someone who compromises my GitHub account (a real risk! But one that I think is well-controlled through 2FA and general account hygiene)
There’s one last secret we have to care about - every job gets a GITHUB_TOKEN automatically created for it. Out of the box for many repos, that token is allowed to push to repos and so a compromised build could itself push content to main that gets picked up by the next build and deploy. We can apply least-privilege here too:
permissions: {}at the top of the workflow, giving it no access to anything by default.contents: readso that jobs that actually do need to read code have read-only access. For us, that is thebuildandverifyjobs, not thedeployjob.persist-credentials: falseonactions/checkoutwhich prevents credentials being left on machines (for later steps in the job / attackers to pick up).
That looks like this, in full:
permissions: {}
jobs:
build:
permissions:
contents: read
# checkout, install Hugo, build, upload artifact
verify:
needs: build
permissions:
contents: read
# checkout, download artifact, check it against the source
deploy:
needs: verify
permissions: {}
environment:
name: production
# download artifact, rsync to the server
(The verify job is a check against what Hugo actually builds which we’ll talk about in a moment.)
Access to the Server#
We’ve done all this work to control access to the secrets but we still have to provide those secrets to the deploy job in order to let it do productive work. To secure the actual access to our prod server, we want to follow the principle of least privilege:
The deploy job logs in as a dedicated
deployuser that has no password, no sudo, and owns nothing but the web root.That user’s key is locked to a single command in
authorized_keys, so it can’t open a shell or run anything else:command="/usr/bin/rrsync -wo -munge /var/www/doomed.ventures/html",restrict ssh-ed25519 AAAA…rrsynconly acceptsrsync, confines it to the web root, and with-womakes it write-only: the key can upload files but never download them.restrictturns off terminals, port forwarding and agent forwarding.The job will only connect to a server that presents the configured host key for the deploy target (
StrictHostKeyChecking=yes).
If this key leaks, the worst an attacker can do is overwrite the site’s files. Which is still bad for me, because they could change my site to advertise a crypto scam or make a false statement.
Malicious Code Execution#
While we’ve insulated the job that actually has e.g. keys to talk to a prod server from potentially-compromised third party packages, there is still a major trust boundary here where potentially-compromised artifacts from Hugo are being deployed to production. Assuming a Hugo breach, it’s easy to imagine that a compromised Hugo binary would produce artifacts that contain some kind of malicious code in their built output.
There are two things we want to try and prevent here:
First, an attacker could compromise Hugo such that it injects some kind of unwanted JS into the page. We can limit some of the damage from this via a well-formed content security policy on the site:
# Everything the site loads is same-origin; nothing is inline.
add_header Content-Security-Policy "default-src 'self'; img-src 'self' data:; object-src 'none'; frame-ancestors 'none'; base-uri 'self'; form-action 'self'" always;
This blocks inline scripts and scripts loaded from other domains, and stops most of the typical exfiltration paths. What it can’t stop is a script file that a compromised Hugo writes into the site itself, since the browser sees that as coming from my own domain. To help us detect that, we can design a verify job that inspects the built artifact for scripts like that - this is actually achievable for my personal site because the only real JS file is the one that powers the site search!
The verify job runs between build and deploy, on its own machine and with no secrets. It downloads the built site and checks out the repo, and then compares the two without ever executing anything from the build:
- The search script is published un-minified, so the job can check that it is byte-for-byte the same file as the one in the repo. Any other
.jsfile in the output blocks the deploy. This is fine, since it’s pretty small and this site doesn’t drive the kind of traffic that would make a difference there. - Every HTML page is parsed and checked against an allow-list of the tags and attributes that the site actually uses. The only
<script>tags allowed are the verified search script and JSON-LD metadata. - Anything that loads or redirects (images, stylesheets, forms,
<meta http-equiv="refresh">) has to stay on my own domain, and every external link has to point to a host that already appears somewhere in the repo’s source.
It has to be a separate job to be worth anything: a check that runs after Hugo on the same machine is running on a machine that Hugo already controls. If everything passes, the job fingerprints the artifact and hands that fingerprint to deploy, which refuses to upload anything that doesn’t match it. This step would let us detect some of the other injections an attacker might want to perform: harmful links or redirects to other sites. This leaves the gap of changed content: words and images.
More creatively, an attacker could hypothetically induce Hugo to create a file that looks innocuous but is actually a symlink (to, for example, /etc/passwd) that they later request once it’s deployed to prod. The upload-artifact action actually replaces symlinks with the content of the actual file - so at worst, this is /etc/passwd on the GHA runner. But this is behavior that has regressed in the past, so it’s worth thinking about. We have a few tools at hand to protect ourselves here.
- Using
rsyncto actually deploy the site to prod without the-lflag results in any non-regular files (such as symlinks) being skipped entirely, so they don’t ever land on the prod server. rrsyncwith the-mungeflag rewrites any paths in a potential symlink to point to a nonexistent (harmless) path. A hypothetical Hugo binary that ships a file callednotes.txtthat is secretly a symlink to/etc/passwd(which the nginx user could read) would instead be rewritten to something like/rsyncd-munged//etc/passwd, which is harmless.- Lastly, the webserver itself (nginx) can be configured to never follow symlinks from the document root with
disable_symlinks on. Even if the first two layers of protection were missing, the symlink would not be followed by the web server and the attacker would just get a 404.
Did We Do a Good Enough Job?#
Assuming that the rest of production server hygiene is being followed (all web services are covered with 2FA, the server itself is locked down in terms of web traffic and SSH, patches are being installed, you know the drill) the remaining risks at this point are going to be:
- A compromise of my accounts: my GitHub account or my hosting account, which I feel pretty good about operationally.
- A malicious release that is pinned by me in good faith. This is an XZ-shaped problem, and the biggest defense here is likely that all dependency updates still depend on me (a relatively lazy human) clicking a button.
GitHub Actions is probably the most widely-used CI system in the world at this point, and so the exercise of working backward from an assumed breach against a real, widely-used target is a fun and useful activity for anyone who is interested in security.