Have you noticed that AWS has regular, significant outages? The web store does as well. The thing is, most of them don't affect the whole store or all customers. Amazon covers for a lot of this with superb customer service (provided you're willing to do it online.)
TL;DR: At the end of writing this I realized I was burying my lead. Since the only way to advance at amazon is to create new stuff, and teams are constantly reformed in the chaos, there are often nobody left to keep past efforts running and things just break down. So at best someone is avoiding the pain of getting yelled at in the weekly site outages, while also being yelled at because they have to constantly churn out new features, all the time.
AWS has some really good solid engineering work in it, no question.
But Amazon is compromised operationally by their completely chaotic engineering practices, which are a result of the political culture and C people in management. (Another person said they knew good team managers, and I am sure this is true in other teams, I just had a bad batch.... but I think as you go up the chain the quality goes down.)
Thus you have people who are between the Sr. Manager and VP level who don't understand technology and are not making decisions based on uptime. Uptime is "important" and thus teams who get blamed for downtime (not always accurately) get pressure on them... but there is a significant lack of anyone empowered to architecturally build a stable system.
They're still using code from the 1990s, and I periodically go to the site and test a couple regression cases I'm aware of that were bugs that were fixed. They've become unfixed in the intervening years, in fact, it appears that the team I left (which had lost %60 of its members at that point, for the same reason I left) was probably disbanded and there simply is nobody in the company who is charged with making sure that this stuff works... or its a responsibility of a team that is forced to spend almost all of its time putting in new features.
TL;DR: At the end of writing this I realized I was burying my lead. Since the only way to advance at amazon is to create new stuff, and teams are constantly reformed in the chaos, there are often nobody left to keep past efforts running and things just break down. So at best someone is avoiding the pain of getting yelled at in the weekly site outages, while also being yelled at because they have to constantly churn out new features, all the time.
AWS has some really good solid engineering work in it, no question.
But Amazon is compromised operationally by their completely chaotic engineering practices, which are a result of the political culture and C people in management. (Another person said they knew good team managers, and I am sure this is true in other teams, I just had a bad batch.... but I think as you go up the chain the quality goes down.)
Thus you have people who are between the Sr. Manager and VP level who don't understand technology and are not making decisions based on uptime. Uptime is "important" and thus teams who get blamed for downtime (not always accurately) get pressure on them... but there is a significant lack of anyone empowered to architecturally build a stable system.
They're still using code from the 1990s, and I periodically go to the site and test a couple regression cases I'm aware of that were bugs that were fixed. They've become unfixed in the intervening years, in fact, it appears that the team I left (which had lost %60 of its members at that point, for the same reason I left) was probably disbanded and there simply is nobody in the company who is charged with making sure that this stuff works... or its a responsibility of a team that is forced to spend almost all of its time putting in new features.