proxysna
2 days ago
Wider audience started to use status pages because the service unreliability became so much more noticeable than before and not the other way around. I never had to use a status page for bear blog or protonmail because i never had and issue with it or just i never noticed.
I am now _required_ to consult status page of github, circleci or MS services etc because i need to know why a build is not passing, why i cannot open a repo, why is my work stalling.
Percentages matter, it is just so much more obvious why they matter when it comes down to important pieces of the internet like github. And i highly doubt the number of 12 hours in the last month. MS has been downplaying the issues they have with GH performance for a while now and i don't think it is time to start to believe them yet. Maintaining these pieces of infrastructure is responsibility and a burden.
Overall i would be careful with "nonlinear significance of numbers near 100%" we are talking gh being well into the 90's this year and one number that infra people are also often being reminded about is that "1% is 3.5 days".
Things are tough for gh people and i feel for them but they are not a startup or a underdog of some sort to receive sympathy in that case.
rdmuser
2 days ago
Yeah I'm at a point where I have a bookmark folder of status pages mostly for very large orgs because I've been using those pages relatively regularly. This is not something I felt the need for historically.
fmbb
2 days ago
Nothing beats downdetector.com anyway. Always quicker!
hinkley
2 days ago
Downdetector also doesn't have a motivation to lie.
I've yet to find a status page that wasn't lying about the actual status.
Also 97% up is bullshit for the 3% of people who are offline.
Saucelabs was doubly bad for this because I'm absolutely certain based on traces that they had some sort of demux bug where they would send events from their tunnel to the wrong job. I could see it in the logs that a test timeout was often the cause of an event firing that was looking for something that never happened, because the event immediately preceding it in the script was never fired. Which meant it was either dropped or went somewhere it shouldn't.
Then it stopped one day and there was nothing in their release notes about it. Lies compounded by further lies.
That's just the most memorable example I have. Stuff like this happens all the time and with many services it plays out the same. There's a perverse incentive not to be transparent about problems with the service, so the status pages play down the intensity of the situation.
Melatonic
20 hours ago
Isnt downdetector just people reporting its down though? Its useful for sure but not actually hooking into any officialy API or anything. Great for when the status page also goes down but surely a lag time
hinkley
19 hours ago
I go to status pages to find out if 1) I’m crazy, 2) if our IT fucked up DNS.
Every service I’ve ever paid for or someone paid for on my behalf has gaslit me about their status page because it’s impolitic and bad for sales to update the page before you know what’s going on, just because some users are reporting issues.
So a third party doesn’t have to deal with VPs kneecapping the engineers’ access to the status page. Or some services can’t update the status page when the site is hard down because they are so obsessed with keeping it up that they have no mitigations when they are down.
I was the one at my biggest gig that had to push to get static 404 and 500 pages uploaded to S3 so we could show something for vanity URLs even if customer ID lookup was down. And then a customer noticed they hadn’t updated since they changed their contact info and I found the job was timing out without an alert or deployment failure for five months. Five. Months. The guy who wrote it had quit, and he didn’t follow my advice on copying a batch job I’d poured way too much effort into. The damned thing was timing out after 50 minutes. I followed my own advice and got it to 4.5 minutes. Almost all of that time delta was waiting for fanout calls, which were pounding the shit out of consumer facing services. 90% of the calls he was making didn’t need to be made.
gplk
2 days ago
Same ! To such a point I ended up installing a menubar app (like vitalsbar). Never felt the need to have a realtime overview to be able to work...
SoftTalker
2 days ago
> 1% is 3.5 days
I think that illustrates the author's point quite well. 99% uptime sounds good, but when you think about a 3+ day outage that doesn't sound very good. Imagine Facebook or TikTok being down for 3 days.
Of course most of the time it's not all one outage, but a bunch of short ones. Still, it might communicate the impact better, especially depending on the argument you're trying to win.
Melatonic
20 hours ago
It used to be that " 5 nines " was the gold standard. I cant believe were now down to 0 for a major service........
fbd_0100
2 days ago
Just two weeks ago protonmail had a massive outage related to data center cooling failure.
nwallin
2 days ago
Microsoft has absolutely gone to shit in the past ~year. Github, Teams, Windows, Azure, Exchange, it's all been fucking trash. Github was running at 80% uptime for a few months, and if they're telling me they're at 98% or whatever now, then they're cooking the books, period. There's no way they're even that reliable.
What's going on at Microsoft? Are they just copy-pasting their github issue reports into copilot and hitting send it without doing code reviews?
fingerlocks
2 days ago
Everyone got laid off. Massive cost cutting across all orgs. Engineers are now evaluated on AI usage and pull request frequency, and not bugs fixed or performance improvements.
VirusNewbie
2 days ago
About four years ago I did interview loops at GCP, Netflix, and Azure at the same time. The latter was a “hiring event” so all my interviewers were from different teams, either managers or TLs.
It was the interview equivalent of the multi-headed dragon meme, where the last one looks absolutely stupid. The contrast was insane, microsoft was an absolute shit show compared to the other two companies in terms of talent, personality, organization and more.
cyanydeez
2 days ago
company just switched to the cloud outlook. Expect it's going to crash several times as microsoft continues to inject it's AI everywhere attitude.