Abstract
Despite the improvements in hardware design parallel systems lack on dependability due to the huge amount of components they consist of. One possibility to introduce fault-tolerance into such systems is backward error recovery where failed modules can be replaced by spares. This work describes an approach to build a fault-tolerant parallel system. Therefore system reconfiguration and recovery based on checkpointing and rollback is presented as well as a fault-tolerant routing algorithm. The enhancement of the acceptance of fault-tolerance is reached by the integration of a user-transparent routing, reconfiguration, checkpointing and rollback protocol. Furthermore, the restriction to a fail-silent failure model (used in many approaches) is released in our work towards a fail-time-bounded behavior.
| Originalsprache | Englisch |
|---|---|
| Zeitschrift | Computer Systems Science and Engineering |
| Jahrgang | 12 |
| Ausgabenummer | 4 |
| Seiten (von - bis) | 245-253 |
| Seitenumfang | 9 |
| ISSN | 0267-6192 |
| Publikationsstatus | Veröffentlicht - 01.07.1997 |
UN SDGs
Dieser Output leistet einen Beitrag zu folgendem(n) Ziel(en) für nachhaltige Entwicklung
-
SDG 9 – Industrie, Innovation und Infrastruktur
Fingerprint
Untersuchen Sie die Forschungsthemen von „Fault-tolerant routing, reconfiguration and backward error recovery for parallel systems“. Zusammen bilden sie einen einzigartigen Fingerprint.Zitieren
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver