🕵️ SicherheitslückenWhat continuous operational resilience looks like under DORA(09.09.2026 um 17:53 Uhr)
🔧 AI Nachrichten OpenAI seeks tougher AI rules. CIOs may feel the ripple effects(10.09.2026 um 12:11 Uhr)
🔧 AI Nachrichten Mistral valued at €21bn after €3bn Series D funding round(08.09.2026 um 10:19 Uhr)
🪟 Windows TippsWindows XP's Cursor Indicator Is Getting a Windows 11 Refresh(25.08.2026 um 13:00 Uhr)
🕵️ SicherheitslückenWhat continuous operational resilience looks like under DORA(09.09.2026 um 17:53 Uhr)
🔧 AI Nachrichten OpenAI seeks tougher AI rules. CIOs may feel the ripple effects(10.09.2026 um 12:11 Uhr)
🔧 AI Nachrichten Mistral valued at €21bn after €3bn Series D funding round(08.09.2026 um 10:19 Uhr)
🪟 Windows TippsWindows XP's Cursor Indicator Is Getting a Windows 11 Refresh(25.08.2026 um 13:00 Uhr)

🔧 Programmierung 🕛 vor 9 Monaten 41 Min Lesezeit
0

AWS re:Invent 2025 - Architecting for hypergrowth: Scaling to 200 million users w/ Skyscanner-ARC209

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht
📺
dev.to

🦄 Making great presentations more accessible.

This project aims to enhances multilingual accessibility and discoverability while maintaining the integrity of original content. Detailed transcriptions and keyframes preserve the nuances and technical insights that make each session compelling.






Overview



📖 AWS re:Invent 2025 - Architecting for hypergrowth: Scaling to 200 million users w/ Skyscanner-ARC209




In this video, AWS and Skyscanner present best practices for scaling applications from initial architecture to serving 200 million users globally. The session covers the build-measure-learn cycle, starting with three-tier web applications using EKS and Aurora DSQL. Skyscanner shares their 10-year journey from a .NET monolith to a cellular EKS architecture running 300 Java services across four regions with 24 production Kubernetes clusters. Key topics include compute options (EC2, ECS, EKS, Lambda), API fronting services (API Gateway, ALB, AppSync), database selection strategies favoring SQL initially, and scaling techniques like CloudFront CDN, ElastiCache, Karpenter for node autoscaling, and multi-region deployments. The presentation emphasizes managing blast radius through cell-based architectures, transitioning to asynchronous communication patterns, cost optimization achieving 95% Spot instance usage, and the importance of observability with CloudWatch. Skyscanner's Flights Live Pricing handles 5,000 searches per second generating 100 billion prices daily, demonstrating practical implementation of these scaling principles.








; This article is entirely auto-generated while preserving the original presentation content as much as possible. Please note that there may be typos or inaccuracies.






Main Part







So let's start from day one, right? We don't have an application yet. Perhaps the only application that I know how to get started with is a three-tier web application, right? That's a front end, a back end, and a data store layer, right? So that's something that we can think about when we first start building our architecture, and we also have this cycle of continuous improvement of building, measuring, and learning. So I'm going to start with my three-tier web application today. Maybe in a year it doesn't exist, it doesn't work out too well. Maybe I change to different types of architecture patterns, we'll see. And then you'll also see too from Skyscanner that they didn't start with their initial architecture on day one, right? So you'll hear later on how Skyscanner has been able to scale over the years, but for now let's start with day one.







So putting it all together, you can think of all of our different compute options as a spectrum. So as you look at the top we have AWS Lambda. You're really only thinking about managing your application code, but as you move further down, think about Fargate, right? You're starting to take on some more of that decision making. So now you're probably thinking about data integrations or security configurations, et cetera. As you keep moving down, you're taking on more of those responsibilities. So where's the best place to start? In my opinion, I think the best place to start is probably around that middle layer because then you can take on additional responsibility. Because if you go lower down you can take on additional responsibility, and that's something to consider when you're thinking about hosting your application.





Now EKS Auto mode fully automates Kubernetes cluster management for compute, storage, and networking. So you can think of this as your personal Kubernetes assistant, optimizing your infrastructure by provisioning optimal compute resources automatically and scaling your node groups based on your workload demands. And it also has multi-AZ for high availability.





So to put it all together, there is no wrong answer for choosing a service, right? So this is kind of our cheat sheet for picking an API fronting service. If you're looking for complex APIs with multiple data sources, go with AppSync. If you need web sockets, throttling, usage tiers, and you're going to scale with millions of requests per month, go with API Gateway. If you have a single API action or method and need billions of requests per day, go with Application Load Balancer.







Now there's probably some of you who are still thinking, I still don't get it. Why should I start with SQL? I'm going to have massive amounts of data. There's no way a SQL database is going to be able to handle this. Now some of you here are probably thinking that's me. I have massive amounts of data. We need to optimize for it today. There's no way a relational database is going to work.





Now why else might you need NoSQL? So NoSQL, again, non-relational database, things like document stores, graph databases, right? But things like super low latency applications is something that you might want to think about for a non-relational database, maybe something like a metadata-driven data set, for example. Something that might exist maybe a couple of years down the line, but maybe from day zero, day one, that may not be something that you encounter quite often.





Again, this isn't most of you if you're starting out from the ground up, so start with SQL databases. So Amazon Aurora is our relational database service. It's durable, which means that your data is stored across three availability zones, and it's fully managed. So you don't have to worry about the underlying infrastructure. You don't have to worry about provisioning or administrative tasks like patching or backups.











Now, again, before we go too much further, we can't tune what we aren't measuring, going back to the cycle of build, measure, and learn. We're kind of missing that measure and learn before we start building again. So we have Amazon CloudWatch for observability. It's built natively on AWS. You can measure things like CPU usage, latency, and request rates. We have something called Real User Monitoring that allows you to collect and view client-side data about your web application.





And using the power of generative AI, we have something called CloudWatch Investigations that allows you to quickly investigate and resolve incidents by surfacing relevant information. So CloudWatch Investigations will take metrics, logs, traces, and other data to generate root cause hypotheses and actionable insights. So now that we have data at hand, we can now make data-driven decisions. So now we're maybe actually seeing slow database queries or slow API requests, things that we can actually address by changing our architecture.





This is one of the ways Skyscanner also scaled their front end. They used CloudFront, which is our content delivery network, and it's built on top of over 700 globally available points of presence. So you want to make sure that you can cache your content closer to your end user to reduce latency.





So let's talk a little bit more about the data tier now that we've addressed the front end. Something that you might be thinking about is going multi-region, right? So when you create a multi-region cluster in DSQL, DSQL actually creates another cluster in a different region and links them together. So these linked regions make sure that you have strongly consistent reads and writes, and then there's a third region called the witness region which basically has limited encrypted transaction logs, and that is used basically to provide durability and availability for these multi-region clusters.





Now the best database queries are the ones you never need to make often, so that's where caching comes into play. With ElastiCache, it's a fully managed service, so you don't have to worry about managing the underlying host. It speeds up reads by storing frequently accessed data in a faster memory location, and Skyscanner actually uses Valkey and Redis clusters to help with that, to help with caching.





Now let's learn a little bit more about the back end here, right? So we can dive a little bit deeper into Kubernetes auto scaling. So there's two main elements to auto scaling in EKS. There's node scaling and pod scaling. So with pod scaling there's horizontal pod scaling and vertical pod scaling. Think of horizontal pod auto scaling as scaling out, right? So you're increasing the number. Think of vertical pod scaling as scaling up, so you're making the pod larger. But if there is no available capacity on the nodes in your cluster, then you want to think about cluster or node auto scaling.





So the Cluster Autoscaler strategy is centered around the use of EC2 auto scaling groups, and Cluster Autoscaler assumes the instance types are identical in a node group. So if you have multiple different types of instances or even like different purchase options like spot versus on demand, you're going to have multiple node groups to support multiple node types, so that really spins up additional clusters making it perhaps difficult to manage, right? And that's where Karpenter comes into play.








Skyscanner's Scaling Context: From a Founder's Loft to 200 Million Users



Thanks, Christine. Hi, hi everyone. So in this talk, I'm going to cover Skyscanner's scaling story over the last 10 years. I'll touch on the scaling context, so what problems that actually drive the scale that we operate on. I'll talk about our 10 year journey of our compute platform. We've had a few day ones in that journey. I'll talk about flights live pricing. That's one of our largest workloads that runs on our platform. And then also talk about some of the cultural and organizational scaling tactics that we use.









Flights metasearch is what drives the scale that we operate at. It looks easy from a UX perspective, like what's difficult about that. But there's some real high cardinality problems in this. So we have billions of unique ticketable flights per year. We have a huge number of partners to integrate. We have multi-petabyte data ingest coming in from these partners, and we see big traffic spikes both from seasonality and from marketing campaigns and other events. So our architecture has to cope with volume, volatility, and variability all at the same time.












Four Generations of Compute: Skyscanner's Evolution to Cellular EKS Architecture



So now I'll talk about scaling compute. Compute is the backbone of our scaling story, so I'll run through four generations of our compute platform. We'll go from hybrid clouds with Auto Scaling Groups and EC2 to large ECS clusters to large Kubernetes clusters, then on to cellular EKS architecture.





So skipping forward to V2, containerization arrived, and by 2018 we had around about 300 services deployed and hundreds of Lambda functions. The services were deployed onto centrally managed ECS clusters. Around about this time, we built our own continuous deployment system called Slingshot, and that allowed us to scale up the number of deployments that we were doing quite significantly. One problem with this architecture though, this was pre-ALB, so every service had an ELB. That proved to be quite costly, so that was one of the downsides of this quite simple approach. Another lesson we learned here is that with a large number of microservices, you can get sprawl. There's a lot of complexity that starts to come into your architecture.







So conceptually we traded mega clusters for a fleet of small composable cells. So why do we use cells? Well, it gives us a guarantee. A failure in one cluster only affects one over N of total capacity. So we cap the cluster size, we can stagger upgrades, and we enforce an N plus two deployment policy for services. So that means the service will be deployed in at least three clusters in the region, sometimes up to five for large services. So this means we can survive multiple cluster failures without dropping availability and without resorting to full regional failover most of the time. So the model here is designed for partial failure as the normal case and not an edge case.





So we also have the scar tissue of operating a cells environment. So 2021, bad config push deleted all application namespaces across 24 clusters in a couple of minutes. It was effectively RM minus RF for Skyscanner. And we published a full write-up if you want to go into the gory details. This incident forced us to get really serious about control plane simplicity and config blast radius and also operational drills.





So the lesson we learned from operating cells environments is the control plane is the most important system you'll build. Don't underinvest in it. We underestimated the migration effort of moving hundreds of services from legacy clusters. It's definitely a marathon, not a sprint. It took us several years to do that. Cells introduced an overprovisioning trade-off, so a small service can be over-replicated and large services can end up in 20 plus clusters, so that's one of the reasons that we're using Spot. And good observability and standardization is critical for preventing your cells architecture from becoming unmanageable.








Flights Live Pricing: Handling 5,000 Searches Per Second with Multi-Layered Infrastructure



So now I'll talk about Flights Live Pricing, because that's one of the largest workloads that we run on the cells environment. So it handles about 5,000 searches per second.



This is the thing that generates the 100 billion prices per day, and we'll see 70 gigabits per minute data transfer at individual regions, so there's a lot going on. If Flights Live Pricing is not available, users can't search for flights, so we have a P1 incident.





We then have an additional layer of Route 53, which uses weighted DNS and health checks to shift traffic between regions. Then within a region, NLB feeds traffic into a cell cluster, and from there Istio handles the last mile routing down to the individual pods. In 2015, we made a very pragmatic decision to run our own NAT instances. This is because of the multi-petabyte data ingest that we have from partners. Managed NAT, although a great service, just wasn't economically viable for us to use.



We do fail over to managed NAT if we have issues with our NAT instances, like we saturate them or there are other problems, but we primarily run on EC2 network-optimized Graviton for all traffic in EU and US. We also use Bring Your Own IP, and so it makes it easier for us to manage IP ranges at our partner APIs. The key point here is you don't have to have managed everything. You can be pragmatic and choose your own pathway through it. And for us, this is about cost control.





Skyscanner is a data business. We emit around about 55 billion events per day into our data lake, which is about 25 petabytes of data under management, and a significant amount of that comes out of Flights Live Pricing. We have regional Slipstream endpoints. That's a custom Go service that runs in our cell's environment, and that writes compressed micro-batches into multiple Kinesis streams. So we learned that not all data is equal, though, and that's why we use different streams with different SLAs and quality profiles.








Cultural and Organizational Scaling: Platform Engineering, Cost Management, and Continuous Improvement



Raw technology alone doesn't scale. You also have to think about abstractions and culture. So in Skyscanner, we have a strong principle of preferring open specifications. That's driven our adoption of Kubernetes, OpenTelemetry, Delta Lake, Istio, Karpenter, and other open source projects. This gives us portable, well-understood abstraction layers at key points in compute, networking, and data.





So we adopted platform engineering before it was cool. This was out of necessity, essentially in 2018. Today we have 40 engineers who operate and evolve our production platform. They cover cloud infrastructure, compute, global traffic routing, observability, CI/CD, and SRE enablement. Roughly 50% of their time is spent operating these systems, another 40% is spent on improving them, like kind of product work, and 10% is on learning and development. So we don't just keep the lights on with our production platform. We're continually moving it forward each day.









Another thing to think about is NoSQL. Christine covered this earlier in a fair bit of detail. And one thing I get a lot of customers asking is when to start looking at NoSQL or DynamoDB in our case. DynamoDB is great for massive scale with very low latency. And some really good use cases are things like key-value data stores, things like metadata. But it's not a one-size-fits-all approach. We've heard a lot this week about AI. AI is going to massively change your data needs. Your data's going to evolve in a massive way as you're doing things. And it's going to give you a lot more data, both structured and unstructured, so put a lot of thought into the database technology you pick and how it's likely to grow with you.





So I mentioned thinking asynchronously. And we've got a diagram here that kind of helps explain it. On the left we have a synchronous command where your client calls Service A, which will then call Service B. If there's any issue with Service B, you won't necessarily get a reply to the client and it'll get held. If we go asynchronous, the client only goes to Service A, it'll give a reply, and it'll separately go to Service B. So say there's an issue between Service A and Service B, maybe a network issue, could be that Service B has some form of issue or latency. The client still gets a good answer.





So transitioning to an asynchronous architecture is an investment that is going to take more time, as you really do need to understand your data and different commonalities with it. Understanding the communication amongst things is really important as well, as are any changes to configuration you need to make. Doing this really gives you a much more in-depth understanding of your application and its architecture though.





So if you're wondering what to use, say you've got a massive throughput of data, you need some ordering, you might have multiple consumers, the ability to replay your data, Kinesis Data Streams is a really good fit. If you're going one to one, you don't have much of a fan out and you're going straight to a target, SNS. If you need an ability to buffer your requests in a queue to have them be consumed, and you can order them or not, SQS is a really good service. And finally, if you've got a one to many fan out with a lot of different targets and schemas, EventBridge is a good option. It's worth saying you can use a number of these in combination. This isn't a one size fits all approach, and there's some really nice integrations between some of these. It will depend on your architecture.








Advanced Scaling Patterns: Cell-Based and Multi-Regional Architectures with Best Practices



Another concept Paul touched upon was the concept of a cell-based architecture. There's something we use quite a lot for our own services at AWS and Amazon. And this is when we deploy a full copy of an application into a cell. The cell is a fault boundary. It's done to reduce the area of impact for any failures. You partition your data, it's effectively sharding, and you have complete isolation between the cells. There



are hard bulkheads between them. You'll have a routing layer at the top which is going to be highly resilient and available. A lot of people look at things like Route 53 for that with its 100% SLA on the data plane. But you will at some point need to be able to scale this. You'll need to know how big they are. So there's a balance you'll have between the management overhead of having cells because you need to manage each one of these, and how much you're willing to have fail in the event of an incident.





Another thing Paul briefly touched upon was multi-regional architecture. Now, a lot of our customers look at multi-regional architectures for different things. That could be for regulatory reasons. They may need to have a certain uptime availability if they're in a regulated industry. They may have another regulatory need to have their data in a particular country, or they may just want to be closer to their customers and have lower latency for them. One key thing if you have multi-region: architect for regional independence. If you have an issue in one region, that shouldn't impact another region. Keep them separate. Try to avoid cross-dependencies. Make your writes idempotent. And like cells, this will create additional overhead for you operationally. Be aware of that before you go in. Really think of the trade-offs you need and what you need to be able to comply with and what your availability needs are. There's a number of talks you'll be able to hear this week and see online around multi-regional architecture, and I'd recommend looking into that if you are looking at that journey.



So I've simplified this with some best practices that apply to both cells and multi-region architectures. Firstly, when you're deploying code and updates, try and use your regions and cells as a way to do it. Deploy very fractionally, very gradually. So only deploy to a small number, a small area, one cell, one region, a small number of users, and if you have an issue, quickly back it out. Reduce that area of impact. It's the way AWS does deployments, very fractionally. Performance test. This is something I see a lot of customers struggle with, but performance testing really is a fantastic way to understand your scale. It'll also let you know at what point you need to then look at moving into another cell or another region, what will start to break, what will start to cause you problems.



Christine touched upon observability earlier, and this is a really critical factor. Make sure it aligns to your fault boundaries. If you've got an issue in one cell, how do you know it's in that cell? Make sure it's clearly identified, tagged, and marked. And finally, when you're going to be testing and observing, look for an outcome. Do you really care too much if an EC2 instance individually has a problem? Probably not. You care about the business outcome. Can your users still find a product? Can they still place an order? Look at those outcomes when you're doing your testing. What's the ultimate experience for your end user?





In closing, we're in a much better place than we even were five years ago with the availability of managed services that will let you out of the box scale much more efficiently. Through the stack, there's a lot more resources available. It's a lot easier for you to do that. The best way to scale is to do less. Christine mentioned the best query you make is the one you don't make. Use caching so you don't actually have to do the work. Let it do the work for you. It's there to get hit. It's there to take that heavy lifting. It will reduce the scope of what your database is actually being queried with and processing.



Refactoring is a big investment. It's a lot of time you're going to have to put into it. It's going to involve some challenges along the way. So make sure when you're doing it, you think carefully about it and the different trade-offs. Look for your best fit technologies based on what you need as you go and be flexible around it. Look at the resilience and availability of your application. We're in a digital world, and people expect to be able to do things when they want, on demand. Any downtime is very damaging to your business, to your brand, and will cause bigger problems. Architect around these fault boundaries. Make sure you understand your area of impact. If you're going to fail, how big is the failure going to be?



So thank you for coming today to our session. Christine, Paul, and I will be available outside for a short time after to answer some questions, and please fill out the session survey in the mobile app. Thank you.






; This article is entirely auto-generated using Amazon Bedrock.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Sam Altman calls GPT-6 Astra rollout ‘messy’ as enterprise users wait for access
1 Quelle
Swiss government explores replacing Microsoft 365 with open-source software
1 Quelle
What continuous operational resilience looks like under DORA
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten AWS re:Invent 2025 - Architecting for hypergrowth: Scaling to 200 million users w/ Skyscanner-ARC209

Thematisch verwandte Begriffe: reInvent, 2025, Architecting, hypergrowth · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...