← Back to home

Distributed Systems Lecture — Research

1. What happens when the problem you want to solve becomes too big for any one computer?

Question

You split the problem up and put these split up pieces to a lot of different computers. Each computer works on its own piece independently at the same time, and then their results are put back together.

Source: research.google — MapReduce: Simplified Data Processing on Large Clusters

2. Suppose 1,000 computers work together. Do you now have one computer that is 1,000 times more powerful? Why or why not?

No, as some parts of a task cannot be split up and must be done in order, which limits the speedup. This is known as Amdahl's Law. For example, if 5% of a task must run sequentially, the maximum speedup is 20 times, no matter how many computers are working together. In addition, computers need to communicate between themselves, which is also slow compared to work done inside a single machine, and every added computer is also another possible point of failure.

Source: wikipedia.org — Amdahl's Law

3. Can 1,000 computers agree on something if some of them fail or even lie?

Yes but within limits. For crash failures, there exist protocols such as Paxos and Raft which remain correct as long as a majority of the nodes are alive, tolerating f failures with 2f + 1 nodes. For “Byzantine” (arbitrary or malicious) failures, tolerating f faulty nodes requires at least 3f + 1 nodes. But no algorithm can guarantee agreement if messages are delayed forever or if even one computer crashes, because a crashed computer looks exactly like a slow one. Real systems work around it by using timeouts.

Sources: raft.github.io — In Search of an Understandable Consensus Algorithm lamport.azurewebsites.net — The Byzantine Generals Problem groups.csail.mit.edu — Impossibility of Distributed Consensus with One Faulty Process

4. When you use ChatGPT, Google, Instagram, or an online game, where is the computation actually happening?

Not on our device. Most of the work happens in remote data centers. For ChatGPT, your prompt goes to big clusters of GPUs that actually run the model. When you search something on Google, it is split across many machines, each searching its own piece of the web index. For Instagram, Meta's data centers store and serve your feed, and your phone is just displaying it. And if you play online games, a game server keeps everyone's actions in sync while your device draws the graphics.

Sources: research.google — Bigtable: A Distributed Storage System for Structured Data arxiv.org — Megatron-LM: Training Multi-Billion Parameter Language Models research.facebook.com — TAO: Facebook's Distributed Data Store for the Social Graph gafferongames.com — What Every Programmer Needs To Know About Game Networking

5. If you could make millions of computers behave like one dependable machine, what could humanity build that we cannot build today?

With that much computational power, we could build much more accurate climate and weather simulations and faster drug discovery by simulating how molecules behave. We could also build extremely secure systems (banking, health records, voting) that never go down and can't be secretly tampered with, and much bigger and more powerful AI models.

Source: wikipedia.org — Fallacies of Distributed Computing

For more information please visit: