RAS Modeling of an HPC Switch System

Dong Tang; W. Bryson; R. Elling

doi:10.1109/PRDC.2008.19

Source

2008 14th IEEE Pacific Rim International Symposium on Dependable Computing > 81 - 86

Abstract

The high end of high performance computing (HPC) systems is now moving toward petascale deployments, delivering petaflops of computational capacity and petabytes of storage capacity. Interconnection of the sheer number of server nodes in an HPC system plays a vital role in the developments. InfiniBand has emerged as a compelling interconnect technology, and provides more scalability and significantly better cost- performance than any other known protocols. This paper presents a reliability, availability, and serviceability (RAS) modeling and analysis of the Sun Datacenter Switch 3456 system, the world's largest standards-based InfiniBand switch, with direct capacity to host up to 3,456 server nodes, against hardware faults. The results show that the system reliability, in terms of connectivity between the server nodes physically connected to the switch, is high for configurations with redundant ports. The study also shows that practicing deferred repair strategies can significantly reduce unscheduled service events and system downtime. Further, the study identifies optimal service strategies by a tradeoff analysis on reliability and availability.

Identifiers

book e-ISBN :	978-0-7695-3448-0
DOI	10.1109/PRDC.2008.19

Keywords

telecommunication switching multiprocessor interconnection networks parallel machines connectivity RAS modeling reliability availability serviceability HPC switch system high performance computing interconnect technology Sun Datacenter Switch 3456 system InfiniBand switch server nodes Maintenance engineering Fabrics Switches Servers Fans Redundancy Markov Model

Additional information

Data set: ieee

Publisher

IEEE

INFONA - science communication portal

RAS Modeling of an HPC Switch System

Source

Abstract

Identifiers

Authors

Dong Tang

Bryson, W.

Elling, R.

Keywords

Additional information

Publisher


Assign to other user
	×
Wrong email address

INFONA - science communication portal

RAS Modeling of an HPC Switch System $("#expandableTitles").expandable();

Source

Abstract

Identifiers

Authors

User assignment

Assignment remove confirmation

You're going to remove this assignment. Are you sure?

Dong Tang

Bryson, W.

Elling, R.

Keywords

Additional information

Publisher

Share

Export to bibliography

Reporting an error / abuse

Sending the report failed

Accessibility options

RAS Modeling of an HPC Switch System