KNOWLEDGE / 10
System Design and Cloud
Distributed system components, architecture trade-offs, and testing cloud solutions.
Questions and practice
Open a question to see the answer, examples, and exercises.
How does RTO differ from RPO?Junior
Answer
RTO is the target recovery time; RPO is the acceptable data-loss window measured in time.
Examples
- An RTO of 30 minutes and RPO of 5 minutes create different service recovery and backup requirements.
Practice exercises
- Propose an RTO and RPO for evidence storage and explain the cost of your choice.
Why does system design start with functional and non-functional requirements?Junior
Answer
Functional requirements define required behaviour, while non-functional requirements define target properties such as latency, throughput, availability, durability, security, and cost. Without numerical assumptions, storage, scaling, or consistency choices cannot be justified. Architecture responds to requirements rather than showcasing fashionable services.
Examples
- A system must accept 500 jobs per second, finish 95% within 30 seconds, and lose no acknowledged jobs.
Practice exercises
- For a file-upload service, define three functional and five measurable non-functional requirements.
Why is a stateless service easier to scale horizontally?Junior
Answer
When an instance keeps no irreplaceable session state locally, any healthy instance can handle the next request. This simplifies load balancing, replacement, and autoscaling. State does not disappear—it moves to shared storage, a cache, or a token, introducing its own consistency and availability trade-offs.
Examples
- A session lives in a shared store, so a user’s requests can reach different instances without sticky routing.
Practice exercises
- Identify local state in a monolithic service and propose ways to externalise it, including risks.
Why use a queue for long test runs?Middle
Answer
It separates task intake from execution and helps manage load and workers; retry and duplicate-handling rules are still required.
Examples
- An API accepts a run quickly while workers consume jobs according to available capacity.
Practice exercises
- Design an Autotest Server with a queue, workers, and redelivery rules.
Which trade-offs does a cache add to a system?Middle
Answer
A cache reduces latency and load on the source of truth but introduces invalidation, stampede, eviction, and consistency problems. Define cache keys, TTL, write policy, miss behaviour, and outage behaviour. A cache should not silently become the only source of critical data.
Examples
- After a price change, write-through updates the cache, while a short TTL limits the impact of missed invalidation.
Practice exercises
- For a product catalogue, compare cache-aside with write-through and design stale-data and cache-outage tests.
How should a consistency model be chosen in a distributed system?Middle
Answer
The choice depends on the business invariant and the acceptable divergence window. Strong consistency simplifies some decisions but may increase latency or reduce availability during a partition. Eventual consistency requires idempotency, conflict handling, reconciliation, and understandable user experience.
Examples
- A view counter may converge later, while charging money twice violates a critical invariant.
Practice exercises
- Classify five product data types by acceptable staleness window and justify the consistency choice.
Why can retries make an outage worse?Senior
Answer
Repeated attempts add load; they need limits, backoff, jitter, and a check that repetition is safe.
Examples
- Thousands of clients retrying immediately can create a retry storm against an overloaded service.
Practice exercises
- Design a retry policy for an unstable API and define its stop conditions.
How does high availability differ from disaster recovery?Senior
Answer
High availability minimises interruption during ordinary component failures through redundancy and failover. Disaster recovery restores the system after large-scale loss of a region, data, or control plane. A replica does not replace a backup when accidental deletion replicates everywhere; recovery plans require regular exercises.
Examples
- A multi-zone deployment survives a node failure, while restoration from an immutable backup covers destruction of the primary database.
Practice exercises
- Design a game day for a region outage: detection, failover, data validation, communication, and return to normal.
How does a capacity model connect performance and cloud cost?Senior
Answer
A model uses arrival rate, concurrency, latency, payload size, storage growth, and headroom to estimate resources. Autoscaling reacts with delay and cannot fix a bottleneck in a database or downstream dependency. Calculate cost per useful unit of work, including idle capacity, network egress, and observability.
Examples
- By Little’s Law, 100 requests per second at 0.2-second average latency imply about 20 concurrent requests before headroom.
Practice exercises
- Estimate compute, storage, and egress for three load levels, then define an autoscaling signal and safety margin.