Here’s a mistake that I made as a programmer and how I recovered.

TIL like attack

The first significant application I built by myself was Hashrocket’s Today I Learned.

Hashocket TIL

A feature of the site was a post “like” button, which you could click to show that you liked a post.

You could even click it when not logged in! The site then kept track of posts you’ve liked. If you’ve ever tried to build a feature like this, you know that it’s hard to constrain unauthenticated user activity.

One day I came into work and found that one of our posts had gained a few thousand likes overnight! We’d been hacked.

Why was this a mistake? “Mistake” might be the wrong word, but it was a shortcoming. We stored a user’s like behavior in browser cookies. And we may have validated those cookies when posting to the likes endpoint (probably, but I don’t remember). What we didn’t do was any other server-side limiting. So, a person could click “like”, open a private browser window, and click “like” again. We thought it wouldn’t happen, but then it did.1

How we fixed it

So, how did we fix it? I’m going to use “we” for many of these actions, but as the principal maintainer, I took many of them and was feeling the pain throughout.

Here are the steps you must take to recover from an event like this:

  1. Start talking
  2. Assess the damage
  3. Fix the bug, the right way
  4. Teach the team

1. Start talking

Speak up and notify those who may be affected. On our small team, saying “Hey, I broke this” was enough.

2. Assess the damage

Second, assess the damage. Here are questions to help:

  • “Who can experience this bug?”
  • “How likely are they to experience this bug?”
  • “How important is that action?”
  • “How long has this bug been live?”

Applied to this problem:

  • “Who can experience this bug?” Everyone who happens to be looking at this post right now.
  • “How likely are they to experience this bug?” Likely, if they’re paying attention.
  • “How important is that action?” Not very; it’s just a social/analytics feature.
  • “How long has this bug been live?” Metadata told us a couple of hours.

It’s important, but not an emergency. Walk, don’t run.

3. Fix the bug, the right way

We’ve bought ourselves some time by being transparent about the issue. As a result, we could now fix the bug the right way.

The first thing we did was disable the feature temporarily:

We also scripted (no console cowboying!) a database reset for that record, so we didn’t lose all the data since the nightly backup. We possibly lost some liking activity on that record during the incident, which we accepted.

Then, we added a configurable server-side rate-limiting module by user IP address, covered by a test:

This feature was offline for a few days, fine for a side project.

4. Teach the team

Teach the team what happened so everybody learns from the mistake.

On our team, my colleague Chris Erin wrote an incident report. Consultancies gain authority from learning in public, so that’s what we did.

Conclusion

Mistakes happen. By following these steps, I learned from it and felt happy with the result.


  1. Getting hacked is no fun. Hindsight has given me a bit of pride that somebody cared enough to hack us. ↩︎