Fault-Tolerant Architecture

#21 · open · 1 comments

View on GitHub ↗

isabelcosta

**Problem** - The basic architecture described above has the disadvantage that, if a broker fails, the entire system stops operating, because events can no longer be routed from publishers to providers. The fault-tolerant variant of the architecture aims at overcoming this limitations. **Objective** - In this version of the system, 3 brokers are assigned to each site. This should ensure that a site remains operational in face of the failure of a single broker at each site. Note 3 that, in the fault-tolerant version, sites are still organised in a tree. However, a broker at a site can now connect to one or more of the 3 servers that may be operational in its parent site. Similarly, a broker may connect to more than one broker at each of its children sites. - Finally, publishers and subscribers may connect to any of the brokers in their site. In case of failures, the system should reconfigure automatically, without disrupting the reliable stream of events. For instance, if the broker to which a publisher is connected crashes, the publisher should automatically reconnect to another broker at its site and subscribers should receive, nevertheless, all the events in the order by which they have been sent. - The students are free to use the replication technique they feel more appropriate to solve the problem at hand. Also, the students are encouraged to use techniques that may leverage from the availability of different brokers to perform some load balancing whenever possible.

Comments

isabelcosta

@vicenterocha @franciscocaixeiro Check Multi-Paxos Algorithm