Settings

Theme

Error 404 (Not Found)

spotify.com

400 points by fairytale · 196 comments (195 loaded)

Reader

82 threads
cronix

back up as of now (10:10a PDT)

This sure makes it easy to know who is hosted by google by going to downdetector.com.

  • judge2020

    Not exactly hard to do that in the first place

      whois $(dig +short spotify.com A)
      
      NetName:        GOOGLE-CLOUD
    • uvdn7

      That only says spotify.com is using GCP's DNS but not necessarily for hosting?

      EDIT: I was wrong, as pointed out by shizcakes.

      • shizcakes

        That's not what that says. It's doing a whois on the hosting IP address, not the DNS service.

      • uvdn7

        While we are at it, looks like spotify.com is using a mix of NS1 and GCP for DNS.

        spotify.com. 172800 IN NS dns1.p07.nsone.net.

        spotify.com. 172800 IN NS dns2.p07.nsone.net.

        spotify.com. 172800 IN NS dns3.p07.nsone.net.

        spotify.com. 172800 IN NS dns4.p07.nsone.net.

        spotify.com. 172800 IN NS ns-cloud-a1.googledomains.com.

        spotify.com. 172800 IN NS ns-cloud-a2.googledomains.com.

        spotify.com. 172800 IN NS ns-cloud-a3.googledomains.com.

        spotify.com. 172800 IN NS ns-cloud-a4.googledomains.com.

        • account42

          How can you name your service NS1 but then not use the ns1. subdomain for the first nameserver. SMH

  • crdrost

    Yeah, per the status page, there's a temporary mitigation rolled out while the team tries to figure out more... so no more 404s but load balancers will be locked down for a while while investigation continues.

  • jonnylangefeld

    Hmm it shows the same spike for instagram and aws as well. Would be funny to me if something on their end depends on GCP. https://downdetector.com/status/aws-amazon-web-services/ https://downdetector.com/status/instagram/

  • fairytaleOP

    Can confirm. https://spotify.com is now working fine.

    • corobo

      Can unconfirm. I can see their website but the app isn't playing nicely at all

      Any non-offline playlists/songs just sit there not playing or telling me I'm offline

      edit @ 37min: App also seems to be working again now

  • technologyvault

    BigCommerce is likely one of those. Our BigCommerce store was down for several hours today.

  • jakub_g

    Github graphql API seems to be throwing error responses for me still

terramex

Etsy is 404 too: https://www.etsy.com

Seems to be a bigger issue.

edit: Nest is down too: http://nest.com

Fitbit.com is 404 too: https://www.fitbit.com

Big GCP issue?

edit2: Downdetector.com shows multiple website and services as down, including Pokemon GO or Rocket League.

GCP status page is still green all over the board: https://status.cloud.google.com

19:10 CET update: Some websites are coming back, including spotify.com, but their app still does not work for me.

information about outage just added to GCP status page, direct link: https://status.cloud.google.com/incidents/6PM5mNd43NbMqjCZ5R...

Description: We are experiencing an issue with Cloud Networking beginning at Tuesday, 2021-11-16 09:53 US/Pacific.

Our engineering team continues to investigate the issue.

We will provide an update by Tuesday, 2021-11-16 10:40 US/Pacific with current details.

We apologize to all who are affected by the disruption.

19:20 CET update:

Description: We believe the issue with Cloud Networking is partially resolved.

Customers will be unable to apply changes to their load balancers until the issue is fully resolved.

We do not have an ETA for full resolution at this point.

We will provide an update by Tuesday, 2021-11-16 11:28 US/Pacific with current details.

Spotify desktop app still not working for me.

19:45 CET: Spotify app is back online for me.

jaredeklee13

Global: Experiencing Issue with Cloud networking

Incident began at 2021-11-16 10:10 (all times are US/Pacific).

https://status.cloud.google.com/incidents/6PM5mNd43NbMqjCZ5R...

mixedbit

Looks like perhaps an issue with Google Load Balancer. We have a load balancer in front of Google Storage Buckets, and can access resources directly from the buckets, but getting 404 when going through the load balancer.

  • te_chris

    Yep, it's the GLB. Went down for us at 5:46 GMT. Just responding 404 and logs reporting an internal error.

  • neom

    Non-engineer here - Is there an easy way to multi-provider redundancy around this? Can you have LBs on multiple clouds and use dns to move around or something? Or does your LB have to be at the provider the app is at? Sorry if this makes no sense. :o

    • sparrc

      yes it's 100% technically possible, the main issue is it would be significantly more expensive.

  • htrp

    I can confirm this on my side too

ksajadi

Half the internet is down because of a Google Cloud global issue on their load balancers, including Spotify and Etsy and GCP status is all green: https://status.cloud.google.com If you ever wondered why GCP is a distant third runner in the enterprise cloud space, here is your answer.

  • yuliyp

    Expecting an instant public post is a bit unrealistic. They had a post up just over 20 minutes after the incident start, which is not that crazy, given that they needed time to triage all of the alarms and understand which component was actually breaking and confirm some technically correct information around it, even if the actual internal incident response can run without the public post.

    • corobo

      It's Google though.. they stop automating everything?

      At least make the screen not show all green or something automatically

      • yuliyp

        Which things should not be green?

        If the automation is working the services will be up. When an incident is happening it's because something is significantly broken, and automation won't properly understand what is and is not working.

        For instance, lots of follow-on alarms might be firing for what are not actually issues with the things being monitored: As an example, I would imagine that datacenter temperatures and fan speeds dropped due to the incident, which might cause automation to suspect a facilities issue, but announcing a facilities issue would be misleading.

        Or metrics around instances live might be tanking as autoscaling groups start downsizing. This would not be an issue with the autoscaling service, and automatically announcing an autoscaling outage would again be misleading.

        In an incident, taking the available data and reaching a conclusion about what is broken and what are effects is something which requires skilled manual effort and is error-prone.

        • corobo

          > Which things should not be green?

          The broken ones is how I usually do it.

          The automation doesn't need to do that, it doesn't need to analyse the situation. It needs to communicate "Hey. Our systems have seen this and have pinged humans, bear with" rather than "nope even though half the internet is down rn, it's all good baby"

          Make a green tick a blue questionmark or something. It doesn't even need to admit fault, it just needs to not be useless. My goal visiting the page is to get a link I can send clients "Updates will be posted here". Nothing more.

          Also if you're hosting your monitoring system on the same system it's monitoring you've just completely missed the point. At least use a different region within your cloud provider, better would be completely different provider. I'd even go as far as using different domains/TLDs to host the page if I was Google sized

        • gowld

          Monitoring should be on a different system, unaffected by an outage in the monitored system.

          • yuliyp

            I think that's tangential to my point. The concerns in my post you replied to about system interdependence making it hard for a monitoring system to separate cause and effect, even if that monitoring system is itself working properly.

  • yupper32

    The big three cloud platforms all have this issue of delayed status updates. Why do you think it's just GCP?

    • mostdataisnice

      ...and there's a reason for that. Automations to update the status page are rarely acceptable, since the status page statuses have legal and financial implications. Therefore, the IM usually has to update it (or tell someone to update it). But, realistically, when you get paged, you first need to figure out what exactly is wrong and at least a vague idea of why. Then, you need to tell someone to update the page. Then, it gets updated.

      The status page will always lag the outage. It's not a conspiracy.

      • deathanatos

        Status pages should be driven that way, though. "legal and financial" implications and "It's not a conspiracy" is a poor excuse.

        Now, I'm on Azure, but it seems like from the comments the situations are similar. So, instead of an automatically updated status page that would help engineers do their jobs, we get a status page that isn't accurate, and customers have pull teeth to get a service credit where/when one is due. And it seems like you can have the cake and eat it too here: while IANAL, a footnote in the SLA or the status page that "this is a machine estimate and not reflective of what goes into the SLA" should do it, no?

      • gowld

        Not updating the status page, to avoid the legal and financial implications, is fraud -- taking money on false pretenses.

        • yuliyp

          fraud? how? what guarantees do they make about timeliness of status updates on their services?

      • kazen44

        Also, in most teams, people who do external communication are different from those doing triage and troubleshooting.

        • cheeze

          Yeah, but they are still people who are responding to a page, working on wording and getting it approved, and then updating.

          20 minutes seems pretty reasonable to me.

  • brown9-2

    AWS typically has the same issue

  • deathanatos

    Azure has the same issues with updating their status page. Sometimes it never happens.

    I might at least hold out some chance that Google Cloud will write an interesting PM, which is something Azure would never do IME.

  • iso1631

    I'm old enough to remember the old claim that the internet was designed to cope with a nuclear attack.

MarcScott

Ironically it appears that IsItDownRightNow? is also down, although that could because they're experiencing what is basically going to be the equivalent of a DDOS.

fsflover

Time to switch to https://bandcamp.com.

  • casi18
    • dymk

      Cool, how do I stream the newest Taylor Swift album on that?

      • benbristow

        You write the domain (ipfs.io) on a blackboard with a chalk pen then take your nails and scratch the board.

        (jk.)

      • iso1631

        My colleague (who loves Taylor Swift) bought the mp3s from somewhere (amazon?) and uploaded them onto her plex server the hour it came out.

        That server continues to work just fine.

        • cheeze

          As a heavy plex user, I can't imagine using it as my default music player. CX isn't great for music, IMO.

cinericius

I wonder if a legal discovery will ever find internal status dashboards that reflect reality rather than fictitious SLA liability-aware status pages.

  • paxys

    You don't need legal discovery for that. Every "X as a service" contract you sign will explicitly say that SLAs aren't dependent on dashboards/ping tests but rather a mostly subjective measure of "availability".

  • hmrr

    Your cynicism is justified and clearly based on the same experience I have :)

algorithm_dk

This is clearly the hottest thing on HN right now, and it was bumped from #1 to #6, anyone knows why? Is it some kind of bot protection?

  • floatingatoll

    User flags, because outages are a fact of everyday life.

    • mbesto

      Which is dumb, linking to status pages shouldn't be on HN. A blog that has analysis and explanations of outages or post mortems should.

      knock-knock dang

      • floatingatoll

        Dang doesn't see messages like that unless you use the footer Contact link, but I remember a comment from him a while back that I would summarize as "Some site users think it's a good use of HN, and other site users disagree and flag it, and we downweight/dedupe them sometimes and/or if someone emails us with the Contact link". I just didn't want you to wait for a reply that'll never come unless you Contact them.

        • mbesto

          hehe I know, was just saying more for fun, but I appreciate the comment none the less.

deforciant

https://linear.app/ is also down

uubk

We found extra rules in our GCLB routing config - removing them restored our service.

al_james

Netlify is also failing for us, and reporting bad TLS certs. Not sure if they use GCP https://www.netlifystatus.com/

jacobkg

We bypassed our Google load balancer and pointed DNS directly at the IP addresses of our servers and that seems to have helped

humanistbot

The title of the 404 page on all the down sites has an extra "1" after the exclamation points: "Error 404 (Not Found)!!1"

humanistbot

Sites that are down according to https://downdetector.com include Spotify, Discord, Snapchat, Etsy, Pokemon Go, Epic Games, Target, Paramount+, Evernote

jaredeklee13

Global: Experiencing Issue with Cloud networking Incident began at 2021-11-16 10:10 (all times are US/Pacific). https://status.cloud.google.com/incidents/6PM5mNd43NbMqjCZ5R...

deberon

GCP outage? Status page shows green but a bunch of sites seem down (Rocket League most importantly).

profmonocle

I wondered why our alerts started going nuts. Seems like basically every global Google Cloud load balancer went down. Doesn't seem to affect single-region network load balancers.

Edit: All of ours are back up. Some other services still seem down though.

soheil

Funny thing is when you google Home Depot or Paramount Plus you get ads served by Google as the first result. When click on it Google then shows you a 404 page. I wonder if they'll get a refund on their Adwords campaign.

  • makecheck

    One of my pet peeves with so many services! Their obnoxious pre-ads can play flawlessly (stealing your time/eyeballs and giving them benefit), and they can still fail to give you the content you exchanged your time/eyeballs to see. Worse, they can repeatedly fail and repeatedly drill the same ads into your brain.

    There ought to be a law that essentially says if ads are “paying” for content, there must be a flawless link between ads and content such that the system can tell if the content is available (or detect after the fact that something was not delivered properly). And then, based on that, it either is required to ensure the ad never plays (since the content cannot be delivered), or that the user must be compensated in some way (e.g. we see you were forced to see an ad but got nothing so we are crediting $1 to your account).

  • progbits

    Why? That's not part of the ad contract. They will get refunded for GCP if it goes out of SLO.

soheil

I also wonder how many companies didn't want to admit they were using Google for their infrastructure. Downdetector shows AWS being affected, it'd be embarrassing if they were caught using Google Cloud Platform.

  • grumple

    Seriously doubt that AWS, Facebook are using Google for infra. There's probably some other effect at play, like people using a Google service to connect to these things. Also don't see any effects on those services personally.

artembugara

ok, so our API is down. We're on GCP...

https://api.newscatcherapi.com/v2/search

collinmanderson

See also https://news.ycombinator.com/item?id=29243740

markbnj

We were down. Just came back. Things seem to be resolving.

IceWreck

https://www.navidrome.org/

Self hosted Spotify. Compatible with subsonic clients.

  • NaughtyShiba

    On one hand, Spotify is much cheaper, on other, perhaps artistd gets paid more (assuming you acquire music legaly)

caffeinated_me

My company was seeing GCP Airflow environments not responding, but they seem to have recovered in the past few minutes.

johanam

https://overleaf.com/ is also 404 now

hs86

https://toggl.com/ is also affected.

dustinmoris

I have a few services running from the same GKE cluster, same ingress controller, same nodes, same GLB, same everything.

Some are 404ing at the moment and others work just fine. Feels like a GLB issue.

Nothing in my GCP dashboard seems to be aware of the issue however.

Only reason I found out is because I use an external service to ping me if a site is down.

trillic

https://www.windy.com down, same issue.

arjan_sch

The 404's changed into 502's.. I guess that's progress. Fingers crossed it's back up soon

Borrible

Yes, but we're still on DEFCON 5.

contrahax

Seeing the same - I have projects in us-east1 that went offline first, then us-west1 went offline a few minutes after. Everything green on their status page and nothing in the dashboard - everything returns a 404 so I'm assuming a really high level LB just took a dump.

  • contrahax

    Seems to have just resolved itself in us-east1 so I'm hoping us-west1 follows a few minutes after.

kadomony

I don’t understand why people post website outages.

Do you think the DevOps teams at these billion dollar streaming companies are so clueless that they don’t have monitoring in place?

Do you think that people who go to a site when it’s down don’t see the same thing?

So whose awareness does this serve?

  • blamazon

    In general, people post things on HN to discuss them. This includes high profile web outages.

mcintyre1994

Looks like it's made it to the Google status page: https://status.cloud.google.com/incidents/6PM5mNd43NbMqjCZ5R...

1cvmask

Is it regional? Surprised Spotify is not active-active on other platforms like AWS and Azure.

dave_aiello

Right now homedepot.com and the APIs that drive their mobile app are down too.

mkl95

Not Google's best week.

iampliny

Might be Google Cloud outage: https://news.ycombinator.com/item?id=29243753

cglace

Everything seems to be working on our end as of 1:08 PM EST.

te_chris

Affecting us. Busiest time of the year and now down 20 min. It's the Global Load Balancer, so god knows what bit of the global edge has been taken out.

nagisa

One of the websites I've noticed this on is back up.

timdaub

- discord doesn't allow me to connect either.

_nickwhite

1:10PM - either it has resolved itself, or a regional issue, but I'm not seeing anything being down from the East coast USA.

lukeschlather

As far as I can tell everything is up it's just that our load balancers aren't routing traffic and just returning 404s.

gassius

Funny enough, datadog, which I was using to investigate on of my vercel sites, is down too

Yeah, Vercel is running some GCP services it seems

  • jeffbee

    Are you able to access your data in the AWS-hosted datadog instance?

    • gassius

      Well, is not like I know how to switch to it, but Datadog came back for me, probably because of that

pdenton

I knew Google would one day take control of the web, just thought they'd have a more clever way of doing it.

DiFronzo

Oh okay, thought something was on my end.

Jansin3

Rocket league perhaps epic games even

scame-miv

I was experienced this issue with my spotify app. Initially thought, it was my internet issue lol.

etimberg

Seeing this across the board with providers on GCP. Firestore however does not appear to be down

pcbro141

https://downdetector.com/

yup

authed

That's why I like the clouds.

cdiddy2

Discord down

Jugurtha

My stuff is running OK on GCP, with GKE usage. Maybe it's related to nameservers?

  • deforciant

    if you are using regional load balancers or serving traffic directly from nodes then you would be fine :) "only" global LB failed

addcninblue

It looks like everything is back now. That was a short outage by recent standards...

thegranderson

Seems like another DNS issue - switching to Cloudflare 1.1.1.1 got me back online...

oussama-gmd

Seems to be an issue with GOOGLE cloud load balancer. our website is down too

Wingy

Seems to be coming back up.

lenniez

Also experience that with all my GCP related infrastructure (Europe)

mey

Our infra in us-central1 behind gcp lb is impacted but not us-west1

davidkuennen

My app was down too. Can confirm it is most likely the GCloud LB

saranshk

Our instances started working again, so seems to be fixed

icecoldfire

Pokémon GO also down

oussama-gmd

Same here, seems to be an issue with load balancers

nicebill8

GCP - my Cloud Run containers are giving 500's

plg

overleaf.com also

crackercrews

CBS.com is down, and some NYT pages as well.

christophclarke

Snapchat also having issues refreshing

melling

I noticed Discord went down for me.

sabbakeynejad

veed.io is down too! Same problem

ishikawa

it was back in a few minutes. But that wasw pretty weird. App Engine was affected.

dustinmoris

Things are back online again!

aalbertson

Lowes search was also down.

dyeje

Discord seems to be down.

ukd1

We're also affected.

NaughtyShiba

Seems to be back already

dfxt8

Discord is down too.

cronix

Seems to be fixed.

milesward

It's back up

Thaxll

It's back!

colewilson

looks like it's back up?

ruined

nice

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection