Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

All of the outages make sense to me as scaling issues. Each month, GitHub is getting the amount of commits they’d normally get in a year. And it’s growing

> Yup, platform activity is surging. There were 1 billion commits in 2025. Now, it's 275 million per week, on pace for 14 billion this year if growth remains linear (spoiler: it won't.) GitHub Actions has grown from 500M minutes/week in 2023 to 1B minutes/week in 2025, and now 2.1B minutes so far this week. So we're pushing incredibly hard on more CPUs, scaling services, and strengthening GitHub’s core features. And as a fine purveyor of hand-crafted shit code for many years, I'm not gonna weigh in on that.

x.com/kdaigle/status/2040164759836778878



That’s no excuse. Particularly if it affects people self hosting their workers.

These are folks that routinely make it a point to press on system design and scalability during interviews.

Now they suddenly can’t scale or design systems but we should accept that?


Even if their actions running servers were at 50% load (or even 20%) at the end of 2025, they'd be screwed right now, at this point it's as much a tech/scaling problem as it is a CapEx problem and a construction one. Given the public's responses to AI specific datacenters you can't pick everything.

Even my own systems at home are burgeoning under the load of the my more ambitious hobby projects I'm doing for fun. I don't envy "we had to build five more datacenters to keep up with demand" class problems just from a how many people have to sign off on them perspective let alone the technical difficulty of doing so


They don't hire well at all. You look at who is on their research team and it's people with social science PHDs, not computer science. Looking at the state of the site it's not a surprise. Highly incompetent people have taken over for a while. Let's not forget a member of the original executive team was a sex pest too.

For like 10 years the only feature development they did was by stealing ideas from GitLab. I wouldn't be shocked if little if none engineering discipline has took place at all during this time if it's this brittle to frequent change.

Guessing it was mostly held together with duct tape and poorly written tests/monitoring systems if any other corporate driven software.

We're really going to find out over the next few years which businesses have good practices or not.


Are you kidding? That's an insane amount of load increase to manage. "Scalability" isn't one thing, especially not at that level, so it's ridiculous to knock them for it. That amount of load ripples across your entire infrastructure.


They have the backing of a trillion dollar corporation and can literally hire some of the best talent on the planet. People nor budgets are an issue here.

Why do we keep giving excuses to poor engineering disciplines + poor management? This problem is entirely GitHub's making and acting like it's some unplanned natural disaster is low key pathetic.


Not hiring 20-25 year olds with social science Ph.D.-s would be a good start (don't remember where I saw that a while ago).

Next: proficient engineering managers, not just people who immediately bend a knee to the middle manager.

Oh wait, I went into sci-fi territory again. My bad.


This was posted elsewhere in the discussion: https://damrnelson.github.io/github-historical-uptime/ Went downhill long before LLMs


The missing line on that graph is November 2019 when Github Actions were introduced. It would also show the point where things went wrong for uptime


Before Github Actions places I worked at pushed monliths to a dedicated CI, Github just hosted the repo, now we push microservices to a monolithic deploy pipeline. I'm sure someday we'll get it right, or maybe the robots will for us.


Back of the envelope math, if anyone wants to correct my mental model: 2.1B (mincore)/week ~ 300M (mincore)/day. Assuming ~300k cores (~3k physical server CPUs, seems fine), that's 1k min/day. Seems about right, though obviously not uniformly distributed.

In any case, the number I'm focusing on here is the 300k cores part (x2 if you're counting in vcores). It does not seem like too much to ask for (significantly) more than that at github scale. It doesn't feel like a hardware issue, is what I'm getting at.


I decided to ask chatgpt for a fermi estimate of cores/datacenter, which you can check for yourself: 1~10 million physical cores. The surprising part of this for me was the "low" power usage (it used 20MW/"datacenter" as another point for estimation).

1MW ~ 6700 NYC citizens' residential usage, apparently, lol. I don't know exactly why, but that citizen number seemed a surprisingly large (well, probably because I've played with approximately-MW lasers).


Your lasers were probably at MW for only a microsecond.


As a comparison, Buildkite is now running 1.5B job mins per week without this downtime.


Anyone have info on if other platforms are seeing the same scaling? And if so, are they seeing problems too? I definitely only hear about GitHub outages, but it is a lot bigger than the others.


Not really a reasonable excuse considering this is completely broken for self hosted runners / paying customers (for a whole day).


But you’re not self-hosting the tasking service


They couldn't handle sending webhooks for eight hours. That seems more than reasonable to expect.


Because github does not offer the option to (and no, running github enterprise doesn't count).


They don't even let you really run those yourself now either, We moved off a big GHE footprint because the support for it was getting abysmal and features were slowly become GH.com. Ive been dreading the move to GitHub.com and it was worse than I expected.


So how's forgejo been going for you?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: